English Language Teaching Educational Journal ISSN 2621-6485 Vol. 8, No. 1, April 2025, pp. 25-36 https://doi.org/10.12928/eltej.v8i1.12747 http://journal2.uad.ac.id/index.php/eltej/index eltej@pbi.uad.ac.id Measuring up: Rasch analysis of English reading comprehension test for informal education learners Arandha May Rachmawatia,1,*, Agus Widyantorob,2 a,b English Language Education Department, Yogyakarta State University, Jl. Colombo No. 1, Karangmalang, Caturtunggal, Depok, Sleman, Yogyakarta, 55281 Indonesia 1arandhamay.2023@student.uny.ac.id *; 2agus_widyantoro@uny.ac.id *corresponding author A R T I C L E I N F O A B ST R ACT Article history Received 9 February 2025 Revised 27 March 2025 Accepted 08 April 2025 This study aims to evaluate the quality of English reading comprehension test instruments used in informal learning, especially as English literacy tests. With a quantitative approach, the analysis was carried out using the Rasch model through the Quest program on 30 multiple-choice questions given to 30 grade IX students from informal educational institutions in Bantul. The results of the analysis showed that although all questions were included in the fit category for the Rasch model, the level of reliability was relatively low, that was 0.52 for items and 0.39 for participants. In addition, 13.3% of the questions showed inconsistent results (misfit), this means that there is inconsistency in the results and quality of the questions that need to be improved. The analysis of the level of difficulty also showed that there were questions that were too easy or too difficult. These findings highlight the importance of revising the test items and the need to increase the number of participants and items to obtain more accurate measurement results. This study also provides practical implications regarding the need for continuous and planned instrument development in the context of informal education, to provide valid and reliable evaluation tools to measure students' literacy skills. © The Authors 2025. Published by Universitas Ahmad Dahlan. This is an open access article under the CC–BY-SA license. Keywords Informal Education Learners Rasch Analysis Reading Comprehension How to Cite: Rachmawati, A.M. & Widyantoro, A. (2025). Measuring up: Rasch analysis of English reading comprehension test for informal education learners. English Language Teaching Educational Journal, 8(1), 25- 36. https://doi.org/10.12928/eltej.v8i1.12747 1. Introduction According to Butterfuss et al. (2020), reading comprehension involves three main elements: the reader, the text, and the reading activity, all of which are situated in a broader social and cultural context. To effectively comprehend a text, readers must have various abilities, such as working memory, inference, attention, motivation, and linguistic and conceptual knowledge. In the context of foreign language learning, Srisanga and Everatt (2021) added that reading comprehension is influenced by low-level skills (such as vocabulary and grammar mastery) and high-level skills (such as making inferences and monitoring comprehension). Both support each other in forming a complete understanding of the contents of the text. In second language (L2) learning, reading comprehension becomes more complicated due to significant differences in experience, institutions, and culture between the first and second languages (Rafatbakhsh & Ahmadi, 2023). Furthermore, test evaluation can also support the development of data-based learning systems that enable personalization of learning according to the reader's ability level (Stenner, 2023). https://doi.org/10.12928/eltej.v8i1.12747 http://journal2.uad.ac.id/index.php/eltej/index mailto:arandhamay.2023@student.uny.ac.id mailto:agus_widyantoro@uny.ac.id http://creativecommons.org/licenses/by-sa/4.0/ https://doi.org/10.12928/eltej.v8i1.12747 26 English Language Teaching Educational Journal ISSN 2621-6485 Vol. 8, No. 1, April 2025, pp. 25-36 Rachmawati & Widyantoro (Measuring up: Rasch analysis of English reading comprehension test) The limited time for learning in formal schools causes the need for additional learning support outside of school. Informal learning makes a significant contribution in the context of language learning because it allows learners to experience authentic, contextual, and continuous learning processes outside the formal classroom. In informal environments, such as everyday social interactions, media consumption, and community involvement, language learners can develop communicative competence naturally and oriented towards meaning. This approach allows them to use the target language in real situations, thereby improving fluency, vocabulary, and cultural understanding (Johnson & Majewska, 2022). In addition, informal learning fosters intrinsic motivation because learners are actively and personally involved in learning experiences that suit their interests and needs (Smith & Seal, 2021). From a critical education perspective, informal learning encourages linguistic empowerment, because learners not only learn language structures but also use them as tools to negotiate meaning and strengthen identities in certain social contexts (Jones & Brady, 2022). Along with the development of technology and digital media, informal spaces such as social media and online platforms also expand opportunities for flexible language learning that is integrated with everyday life (Smith, 2021). Thus, informal learning in language learning offers a holistic, relevant, and experience-based approach, which enriches learners' linguistic and socio-cultural competencies. The concept of language learning outside the classroom or informal context has attracted great interest among educators and researchers (Chong & Reinders, 2022). There are still limited studies evaluating the quality of test instruments in the context of informal education, especially in the Bantul area. Several previous studies have shown the effectiveness of the Rasch model in evaluating and validating various types of language tests, including reading comprehension tests. Aryadoust et al. (2021) conducted a comprehensive review of the application of the Rasch model in language assessment and emphasized the importance of reporting unidimensionality, local independence, and item and person reliability. Dunn (2024) developed the Rasch model with a Generalized Linear Mixed Model approach to analyze lexical factors that influence the level of item difficulty in vocabulary tests. Meanwhile, Morea et al. (2024) used Rasch analysis to validate a receptive vocabulary test in young language learners, showing the importance of the match between item difficulty and test taker ability. Nguyen (2022) combined Rasch analysis and CFA to validate the construct structure of a college student reading achievement test, and found that a one-factor model was the most appropriate to represent reading ability. Anggia and Habók (2023) highlighted the importance of adjusting text complexity and task difficulty in developing a Rasch-based reading test for EFL students, which can improve the fit between reader ability and text characteristics. A study by Noroozi and Karami (2024) also confirmed the relevance of the Rasch model in evaluating various aspects of validity based on the Messick framework in a university-level English proficiency test. Meanwhile, Polat (2022) and Toker and Seidel (2023) applied variants of the Rasch model such as the Many-Facet Rasch Model and the Mixture Rasch Model to evaluate rater behavior and respondent heterogeneity, highlighting the flexibility of this model in the context of complex language assessment. Educational strategies are needed to develop students' abilities in informative learning by organizing methods and reflecting on the environment as a media space for transferring information (Tkacová et al., 2022). The concept of Rasch Analysis refers to an approach in Item Response Theory (IRT) that is used to measure individual abilities and the quality of test items objectively and accurately. This model was developed by Georg Rasch and is based on the probabilistic principle, where the probability of a person answering an item correctly is predicted by the difference between the person's ability and the item's difficulty level. Rasch analysis assumes that good data is data that fits the model, not a model that is adjusted to the data (Subagja et al., 2023; Winarti & Mubarak, 2020). Technically, Rasch analysis allows the transformation of ordinal scores into interval scales (logit), allowing direct comparison between respondent ability and item difficulty on a single measuring line. Thus, Rasch provides a reliable and fair way to evaluate individual ability and item quality without relying on sample characteristics or score distributions (Rizbudiani et al., 2021; Priyani & Sugiharto, 2024). In addition, Rasch analysis allows for examination of item characteristics such as statistical fit (infit and outfit MNSQ), unidimensionality, local independence, and Differential Item Functioning (DIF), which are useful for detecting bias between respondent groups (Christensen & Ammentorp, 2024; Prabowo & Rahmadian, 2023). According to Luber et al., (2020), although in the RASCH model, item selection for developing tests is a problem, this is an effort to plan so that the quality of the test meets needs and objective. ISSN 2621-6485 English Language Teaching Educational Journal 27 Vol. 8, No. 1, April 2025, pp. 25-36 Rachmawati & Widyantoro (Measuring up: Rasch analysis of English reading comprehension test) This study aims to evaluate the quality of English reading comprehension test in the context of informal institutions through the Rasch model analysis. This study also provides the empirical insights into reliability, the suitability of test items and participants, and the level of difficulty of the questions used. Furthermore, it enables to encourage the test instrument to be better for the informal learning context in the future. 2. Method 3.1. Research Design This research is quantitative research that uses an approach to evaluate objective theory by testing the relationship between variables that can be measured using research instruments. The data obtained can be analyzed using statistical procedures. This research was designed as survey research, which focuses on collecting information and research samples through questionnaires as the main tool for collecting data and aims to obtain an overview of various aspects of the population group. The modeling used in this research is item response theory (IRT). This is a probabilistic model that attempts to explain the relationship between each test taker's response to the test items through the ability variable, which is measured using an instrument. 2.2 Participants The participants for this research classified in Table 1. were 30 students of class IX of Junior High School from one of the informal education institutions in Bantul. The sampling technique uses convenience sampling, namely selecting participants who are willing to be samples in the research (Creswell, 2014). In an effort to minimize bias and errors in responses, there is a briefing before the test was carried out. Table 1. Data of Participants Students’ Gender Grade Total Female IX 20 Male IX 10 The data collection process was carried out using Google Forms from 15 to 22 December 2024. Before asking test participants to fill in answers on the test instrument, test participants are first given an understanding of the procedures and confidentiality of the data provided. Furthermore, test participants are also given an explanation regarding the material being tested so that test participants can more easily fill in the answers on the instrument. 2.3 Instrument This research focuses on developing questions related to reading comprehension as English literacy test. The reading used in this test instrument includes articles, reports, textbooks, etc. (Brown & Abeywickrama, 2018). Based on the objectives of the test instrument, the researchers used descriptive, narrative, greeting, advertisement, recount, and report texts. These types of text were chosen because they fit the context of the English materials that fit. on Merdeka Curriculum standards. The most of the questions used in this instrument were adapted from previous tests to balance the level of differences based on the Bloom Taxonomy. This foundation is used to guide variations at the cognitive level, remembering, understanding, and analyzing text. Details and specifications of questions on the test instrument with a total of 30 multiple choice questions have been presented and sorted by type in Table2. Table 2. Test Item Specification Questions Type Number of Questions Item Numbers Identifying information from a text 2 4,25 Understanding information in a text 2 1,7 Understanding the meaning of words or phrases in context 4 2,5,8,15 Understanding the main idea of a text 1 6,21 28 English Language Teaching Educational Journal ISSN 2621-6485 Vol. 8, No. 1, April 2025, pp. 25-36 Rachmawati & Widyantoro (Measuring up: Rasch analysis of English reading comprehension test) Questions Type Number of Questions Item Numbers Interpreting information from infographic data 2 26,9 Concluding which readers benefit from the text 3 11, 12, 13 Applying the correct conjunction to complete a gap sentence 1 20 Analyzing the purpose of a text 4 3,22,27,29 Analyzing correct information based on the text 3 10, 19,28 Analyzing the main information from a text 2 16,24 Analyzing the possible outcome of a specific action based on the text 2 17,23 Identifying relationships between ideas and predicting the main idea of a missing paragraph 2 18,30 Predicting actions or responses of the reader 1 14 Total Questions 30 2.4 Data Analysis Data analysis in this study used the Rasch model approach from the Item Response Theory with the help of the Quest program. This model was chosen because it is able to provide detailed information regarding the quality of the test items and the abilities of the test participants, as well as separating instrument parameters from participant characteristics. There are four main components in this analysis, namely reliability, item and participant suitability, question difficulty level, and identification of unclear items. First, reliability analysis was conducted to evaluate the consistency of the instrument (item reliability) and the consistency of participants' answers (case reliability). Reliability values were analyzed based on Quest output and interpreted using categories commonly used in Rasch analysis, such as high, medium, and low. Second, the fit of the items and participants to the Rasch model was analyzed using two parameters, namely Infit Mean Square (MNSQ) and Infit t. The range of Infit MNSQ of 0.77 to 1.33 was used as the fit limit, according to the standard of Wright and Mok (2004), which indicates that the participant's response is still within the acceptable variation limits. In addition, the Infit t value with a range of -2 to 2 is used to identify participant responses that deviate statistically from model expectations. Third, the level of difficulty of each question item is analyzed based on the threshold value in the logit generated from the Quest program. The logit value is then classified into five categories, namely very easy (b < -2), easy (-2 ≤ b < -1), moderate (-1 ≤ b ≤ 1), difficult (1 < b ≤ 2), and very difficult (b > 2). This grouping helps in assessing the balance of the distribution of the level of difficulty of the questions in the instrument. Finally, identification of problematic questions is done by marking questions that have a perfect score value, namely questions that are answered completely correctly or completely incorrectly by the participants. Questions like this are considered not to provide meaningful information and are categorized as misfits in the Rasch model, so they need to be revised or eliminated. The results of this entire analysis process are then interpreted by referring to related literature to determine the overall quality of the instrument, as well as to provide recommendations for the use, revision, or elimination of certain items in the context of measuring English literacy in informal education. 3.5. Validity and Reliability The validation process in research is a process that refers to the accuracy or truth of research tools and interpretations taken through research findings (William, 2024). The validation of the instrument was conducted by an English lecturer and a teacher as the expert judgments who inditifiy the content of items in the instrument. Furthermore, the validity and reliability process will be presented based on the result of classical test theory (CTT) and item response theory (IRT). This tool also has important features as a measure of item and person reliability, item suitability to the model, and the level of difficulty of each item (Faradillah & Febriani, 2021). Thus, the validity and reliability of this research instrument can be seen in the results to ensure the quality of the scores on the tests being tested. ISSN 2621-6485 English Language Teaching Educational Journal 29 Vol. 8, No. 1, April 2025, pp. 25-36 Rachmawati & Widyantoro (Measuring up: Rasch analysis of English reading comprehension test) 3. Findings and Discussion 3.1. Test Reliability The reliability analysis of the instrument was conducted using the Rasch model with the help of the Quest program. The results of the test reliability analysis on instrument items based on the RASCH model in this study are presented in Table 3. Table 3. Statistical Summary of Item and Person Estimates Estimates Mean SD SD (adj) Reliability Infit Mean Square Outfit Mean Square Infit t Outfit t Mean SD Mean SD Mean SD Mean SD Item .00 .93 .67 .52 1.00 .15 1.19 .74 -.01 .65 .20 .87 Case 1.94 .78 .49 .39 1.00 .25 1.19 .77 -.01 .85 .18 .90 Table 4. Item and Person Reliability in Rasch Model Fit Indicates Interpretation Infit MNSQ Interpretation Outfit t Interpretation <0.67 Low >1.33 Misfit ≤ 2.00 Fit 0.67-0.80 Sufficient 0.77-1.33 Fit ≥ 2.00 Misfit 0.81-0.90 Good <0.77 Misfit 0.91-0.94 Very Good >0.94 Excellent Table 3 presents a summary of the statistics of the item and participant estimates in the Rasch model, focusing on the aspects of reliability and data fit. From the table, it is known that the item reliability value is 0.52, while the participant reliability is 0.39. According to the standards commonly used in measurement theory, this reliability value is relatively low, both for the instrument and the respondents. Low reliability indicates that the instrument is not consistent enough in measuring the participants' abilities, and there is a possibility of large score fluctuations if the test is re-administered under similar conditions. One of the main causes of this low reliability is the limited number of participants and items, namely only 30 students and 30 questions. However, when reviewed from the Infit Mean Square (MNSQ) and Outfit t values, it was found that all items were within the acceptable range of values. The average Infit MNSQ value for items was 1.00 with a standard deviation of 0.15, and for participants it was also 1.00 with a standard deviation of 0.25. This value is in accordance with the criteria in Table 4, which states that an item is categorized as fit if it has an Infit MNSQ value between 0.77 and 1.33. The Outfit t value on the item also showed supportive results, with an average of 0.20, within the fit range ≤ 2.00. The Infit t for the item was -0.01, this means that the participants were also classified as very close to ideal score. The suitability of the test items to the Rasch model as seen from the Infit and Outfit values shows that even though the reliability value is low, the instrument is still statistically appropriate to the assumptions of the Rasch model. This reflects the nature of the Rasch model which places more emphasis on the suitability of the response pattern to the mathematical model, rather than relying solely on the consistency of the results (reliability). In other words, the Rasch model can provide useful information about the performance of items and participants even when reliability is low. However, reliability remains an important aspect in instrument evaluation. Low reliability values may indicate that the instrument has not been able to distinguish participants' abilities sharply. Therefore, increasing the number of questions and participants is highly recommended so that the instrument produces more stable and reliable measurements. That way, both aspects of suitability and consistency can be achieved simultaneously to improve the quality of the test instrument. 30 English Language Teaching Educational Journal ISSN 2621-6485 Vol. 8, No. 1, April 2025, pp. 25-36 Rachmawati & Widyantoro (Measuring up: Rasch analysis of English reading comprehension test) 3.2. Estimation of Person Fit After discussing the results of the reliability tests on all items and cases, the next stage is related to the person fitting into the RASCH model. The purpose of these estimates is to provide supporting or contradictory findings. Table 5. Person Fitting the Rasch Model Items Infit MNSQ Infit t Criterion Items Infit MNSQ Infit t Criterion 1 1.06 .41 Fit 16 .93 -.19 Fit 2 Perfect score Misfit 17 1.12 .39 Fit 3 1.38 1.06 Fit 18 .97 .14 Fit 4 .90 -.16 Fit 19 Perfect score Misfit 5 1.04 .22 Fit 20 .75 -1.10 Fit 6 .99 .13 Fit 21 1.43 1.54 Fit 7 .65 -1.76 Fit 22 .72 -.72 Fit 8 .81 -.31 Fit 23 .83 -.55 Fit 9 .71 -.91 Fit 24 .65 -1.32 Fit 10 .77 -.88 Fit 25 1.13 .43 Fit 11 1.32 1.03 Fit 26 1.26 .70 Fit 12 1.18 .54 Fit 27 1.25 .60 Fit 13 .65 -1.32 Fit 28 Perfect score Misfit 14 1.31 .68 Fit 29 1.29 .96 Fit 15 .97 .05 Fit 30 Perfect score Misfit Criterion for fit person: 0. 77 ≤ Infit MNSQ ≤ 1, 33 OR -2 ≤ Infit t ≤ 2 Table 5 presents the results of the person fit estimation analysis, which aims to determine whether the answer patterns of test participants are in accordance with the expectations of the Rasch model. In this model, a person is categorized as fit if the Infit Mean Square (MNSQ) value is in the range of 0.77 to 1.33 or the Infit t value is in the range of -2 to 2. The results of the analysis show that out of 30 test participants, 26 participants (86.7%) were in accordance with the fit criteria, while 4 participants (13.3%) were categorized as misfit. Fit participants showed that their responses to test items were consistent with the predictions of the Rasch model, meaning that they answered easy items correctly and were more likely to answer difficult items incorrectly, as expected. This reflects that most participants followed a rational response pattern appropriate to their ability level. In contrast, four participants who were classified as misfit had answer patterns that deviated from the model predictions. Three of them obtained a perfect score, that is, answering all questions correctly. Although at first glance it looks positive, in the context of Rasch, a perfect score actually complicates the analysis because it cannot show variations in ability to answer questions with different levels of difficulty. One other participant showed an Infit MNSQ or Infit t value outside the criterion limit, which was likely caused by random answers, unintentional, or lack of concentration during the test. This issue is essential to observe because although in general the majority of participants are fit; the presence of misfit participants can affect the accuracy of the ability parameter estimation. However, because the number is only 13.3%, its influence on the overall analysis results is relatively small. In the context of instrument development, these results indicate that in general participants can understand and respond to questions consistently, and the instrument is quite good at measuring participants' abilities according to the assumptions of the Rasch model. Overall, this person fit analysis strengthens the validity of the instrument, because the most participants show consistency in their answer forms. ISSN 2621-6485 English Language Teaching Educational Journal 31 Vol. 8, No. 1, April 2025, pp. 25-36 Rachmawati & Widyantoro (Measuring up: Rasch analysis of English reading comprehension test) 3.3. Estimation of Item Fit The other part is the estimation of item fit. This aims to show whether the test items on the instrument have functioned well or not based on the person-fit criteria in the previous section. Items fitting are also estimated individually using the results of the QUEST program output. The results of the output show a collection of information related to the suitability and feasibility of each test item for the RASCH model. The results of the items fitting analysis in this research are illustrated in Table 6. Table 6. Items Fitting the Rach Model Items Infit MNSQ Infit t Criterion Items Infit MNSQ Infit t Criterion 1 .98 .00 Fit 16 .85 -1.00 Fit 2 .87 -.20 Fit 17 .88 .00 Fit 3 1.03 .20 Fit 18 1.11 .40 Fit 4 .85 -.10 Fit 19 1.18 .50 Fit 5 .75 -1.50 Fit 20 .99 .10 Fit 6 .75 -1.50 Fit 21 .82 -1.50 Fit 7 1.29 .90 Fit 22 1.18 .50 Fit 8 1.00 .30 Fit 23 Perfect score Misfit 9 Perfect score Misfit 24 1.08 .40 Fit 10 1.18 1.10 Fit 25 Perfect score Misfit 11 .95 -.30 Fit 26 .83 -1.30 Fit 12 .79 -.40 Fit 27 .82 -.30 Fit 13 .97 .10 Fit 28 1.06 .30 Fit 14 .91 -.10 Fit 29 1.17 .50 Fit 15 1.07 .30 Fit 30 1.10 .40 Fit Criterion for fit items: 0,77 ≤ Infit MNSQ ≤ 1,33 OR -2 ≤ Infit t ≤ 2 Table 6 presents the results of the item fit analysis based on two main indicators in the Rasch model, namely Infit Mean Square (MNSQ) and Infit t. These two indicators are used to assess the extent to which each item functions as it should in measuring the participants' abilities. In the context of the Rasch model, an item is considered fit if the Infit MNSQ value is between 0.77 and 1.33 or the Infit t value is in the range of -2 to 2. The results of the analysis showed that out of 30 questions, 27 (90%) were included in the fit category, while 3 (10%) were categorized as misfit. The three misfit items were questions number 9, 23, and 25, all of which received perfect scores from the participants. This means that all participants answered the questions correctly, so that there was no variation in responses that could be analyzed. In the Rasch model, this issue is considered uninformative because the questions are unable to differentiate abilities between participants. These types of questions are considered too easy (or in some cases too difficult), so they do not contribute to effective measurement. Meanwhile, the other 27 items encountered the suitability criteria. The Infit MNSQ values on these items were within the specified tolerance limits, indicating that most of the questions had functioned optimally in measuring participants' reading ability. This indicates that the questions were able to provide valid and representative information regarding the variation in participants' abilities based on the level of difficulty of each item. These findings generally indicate that the quality of the items in this instrument is quite good, because most of them meet the statistical requirements in the Rasch model. However, the existence of misfit items should be a concern, because it can reduce the effectiveness of the instrument as a whole. Therefore, items that show perfect scores should be revised or replaced with items that have a moderate level of difficulty in order to provide more informative response variations and improve the quality of measurement. 32 English Language Teaching Educational Journal ISSN 2621-6485 Vol. 8, No. 1, April 2025, pp. 25-36 Rachmawati & Widyantoro (Measuring up: Rasch analysis of English reading comprehension test) 3.4. Item Difficulty Level To find out the level of difficulty of an item through the QUEST program, you can find out by looking at the data from the item estimate (Threshold) analysis. In Table 7, the data resulting from the analysis is described with the criterion for each item. Table 7. Items’ Difficulty Level Items Threshold Interpretation Items Threshold Interpretation 1 .58 Moderate 16 1.17 Difficult 2 -.28 Moderate 17 -.74 Moderate 3 -.28 Moderate 18 -1.47 Moderate 4 -.74 Moderate 19 -.74 Moderate 5 -.28 Moderate 20 -.24 Moderate 6 1.19 Difficult 21 1.08 Difficult 7 .06 Moderate 22 -.74 Moderate 8 .34 Moderate 23 Missing Very Easy 9 Missing Very Easy 24 -1.47 Easy 10 1.19 Difficult 25 Missing Very Easy 11 1.19 Difficult 26 2.05 Very Difficult 12 -.28 Moderate 27 -.28 Moderate 13 -.28 Moderate 28 -.74 Moderate 14 .06 Moderate 29 -.28 Moderate 15 -.76 Moderate 30 -.28 Moderate Criterion b>2 Very difficult -1≤b≤1 Moderate -1≤ b ≥-2 Easy 1