{"id":246,"date":"2026-09-07T09:19:47","date_gmt":"2026-09-07T07:19:47","guid":{"rendered":"https:\/\/iq-test.org\/blog\/?p=246"},"modified":"2026-09-07T09:19:49","modified_gmt":"2026-09-07T07:19:49","slug":"iq-test-reliability","status":"publish","type":"post","link":"https:\/\/iq-test.org\/blog\/en\/iq-test-reliability\/","title":{"rendered":"43,651 Tests Analyzed: How Reliable Is the IQ Test?"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\"><strong>In short:<\/strong> I analyzed 43,651 historical test sessions from IQ-Test.org. The result is strong: Taken together, the 46 questions included in the analysis show remarkably consistent performance. Cronbach&#8217;s alpha and McDonald&#8217;s omega are both approximately 0.88. Most individual questions are also meaningfully related to the overall score, and the response patterns reveal a clear common performance factor.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">However, this internal analysis does not show whether the reported IQ score would exactly match the result of a professionally administered intelligence test. It first examines how reliably the test works as a measurement instrument in its own right. Reliability, validity and other requirements for a sound test are related, but they are not the same thing. [1]<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Online IQ tests often use words such as \u201cscientific,\u201d \u201caccurate\u201d and \u201creliable.\u201d These claims are only useful when it is clear what was actually examined.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I have been running IQ-Test.org for more than ten years. During that time, the site has accumulated a large body of test data, allowing me to examine the test much more closely today than when it was first launched. I focused on four questions:<\/p>\n\n\n\n<ol start=\"1\" class=\"wp-block-list\">\n<li>How consistently do the questions work together as one test?<\/li>\n\n\n\n<li>How easy or difficult are the individual questions?<\/li>\n\n\n\n<li>How well do they distinguish between lower and higher overall performance?<\/li>\n\n\n\n<li>What common structure can be found in the response patterns?<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">The main part of this article explains the results in accessible language. Readers who want to examine the calculations more closely will find a detailed methodology section and the scientific sources at the end.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Key results at a glance<\/h2>\n\n\n\n<figure class=\"wp-block-table is-style-stripes\"><table class=\"has-fixed-layout\"><tbody><tr><th>Measure<\/th><th class=\"has-text-align-right\" data-align=\"right\">Result<\/th><th>What it means<\/th><\/tr><tr><td>Test sessions analyzed<\/td><td class=\"has-text-align-right\" data-align=\"right\">43,651<\/td><td>The analysis included historical cases that passed internal checks and were marked as suitable for norming.<\/td><\/tr><tr><td>Questions analyzed<\/td><td class=\"has-text-align-right\" data-align=\"right\">46<\/td><td>These questions had been scored consistently throughout the period covered by the analysis.<\/td><\/tr><tr><td>Cronbach&#8217;s alpha<\/td><td class=\"has-text-align-right\" data-align=\"right\">0.876<\/td><td>The questions show high internal consistency.<\/td><\/tr><tr><td>McDonald&#8217;s omega<\/td><td class=\"has-text-align-right\" data-align=\"right\">0.875<\/td><td>A different statistical model produces virtually the same result.<\/td><\/tr><tr><td>Questions with corrected item-total correlation \u2265 0.30<\/td><td class=\"has-text-align-right\" data-align=\"right\">32 of 46<\/td><td>Around 70% of the questions are clearly related to performance on the rest of the test.<\/td><\/tr><tr><td>First factor<\/td><td class=\"has-text-align-right\" data-align=\"right\">Eigenvalue 12.98<\/td><td>The response patterns reveal a dominant common performance factor.<\/td><\/tr><tr><td>Loadings on the first factor \u2265 0.30<\/td><td class=\"has-text-align-right\" data-align=\"right\">39 of 46<\/td><td>Most questions make a recognizable contribution to this common factor.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">What data did I analyze?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">I included 43,651 historical test sessions that had passed an internal plausibility check and were marked <code>norm_valid = 1<\/code>, meaning they were considered suitable for use in the norming dataset. For this analysis, I examined only the answers to the test questions. Age, gender and other demographic characteristics were not included.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Each question was scored as either correct or incorrect. Every completed test therefore produced a sequence of zeros and ones: 1 for a correct answer and 0 for an incorrect answer. The total score used in this analysis was the number of correct answers across the 46 usable questions.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">On average, participants answered 24.45 of the 46 questions correctly. The median was 25, and the standard deviation was 8.00 points. The scores therefore varied enough to examine differences between participants and how the individual questions performed.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Why I analyzed 46 rather than 47 questions<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">When I reviewed the historical data, I found a problem with question 40 (q40): For a period of time, the incorrect answer key <code>115<\/code> had been stored for this question. The database records only whether an answer was considered correct or incorrect according to the key in use at the time. The participants&#8216; original answers are no longer available.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That means the historical data cannot be genuinely corrected after the fact. I therefore excluded q40 from the total score, reliability calculations, item-total correlations and factor analysis. All 43,651 test sessions could still be retained in the sample.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Question 40 has since been corrected in the live test. Once enough new responses have been collected and scored consistently, I will be able to analyze it again.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">How reliably do the questions work together?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">In this article, \u201creliability\u201d primarily refers to the <strong>internal consistency<\/strong> of the overall scale: Do people who solve many questions also tend to perform well on other questions in the same test?<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">An internally consistent test contains questions that contribute meaningfully to a shared overall result rather than behaving as unrelated or contradictory tasks. The questions do not have to be identical. What matters is that they work together in a coherent way.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Cronbach&#8217;s alpha: 0.876<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Cronbach&#8217;s alpha is a widely used measure of internal consistency. Put simply, it considers how strongly the questions are related to one another in relation to the variation in total test scores. The coefficient was introduced by Cronbach in 1951. [2]<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For the 46 questions analyzed here, alpha is <strong>0.876<\/strong>. This is clear evidence of high internal consistency for the overall scale.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">On its own, however, alpha does not prove that the test measures only one ability or that the reported IQ score is correctly calibrated against an external standard. It is also influenced by the number of questions in the test. I therefore interpret alpha together with omega, the corrected item-total correlations and the factor analysis. [4]<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">McDonald&#8217;s omega: 0.875<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">McDonald&#8217;s omega addresses a similar question but is based on a factor model. Unlike alpha, this model can account for the fact that individual questions may be related to the common measurement target to different degrees. [6]<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">At <strong>0.875<\/strong>, omega is almost identical to alpha. Two measures built on different assumptions therefore lead to the same conclusion: The total score across the 46 questions has high internal consistency.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">How difficult are the individual questions?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">For questions scored as correct or incorrect, <strong>item difficulty<\/strong> is straightforward: It is the proportion of participants who answered a question correctly.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">If 80% of participants solve a question, its value is 0.80\u2014or 80%. Under this definition, a higher percentage means an easier question in this sample.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"576\" src=\"https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-item-difficulty-43651-1024x576.png\" alt=\"Percentage of correct answers for each usable question.\" class=\"wp-image-250\" srcset=\"https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-item-difficulty-43651-1024x576.png 1024w, https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-item-difficulty-43651-300x169.png 300w, https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-item-difficulty-43651-768x432.png 768w, https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-item-difficulty-43651-1536x864.png 1536w, https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-item-difficulty-43651.png 1600w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\"><em>Figure 1: Percentage of correct answers for each usable question. Based on my analysis of 43,651 test sessions. Some information has been obscured to protect the test; further details are available on request.<\/em><\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The proportion of correct answers ranges from <strong>98.13% for q2<\/strong> to <strong>2.19% for q46<\/strong>. The median across all questions is <strong>53.76%<\/strong>, placing the typical question near the middle of the observed difficulty range.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Overall, 37 of the 46 questions were answered correctly by between 20% and 80% of participants. Four questions were solved by more than 90%, while two were solved by no more than 10%.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This wide range of difficulty is a strength of the test. Easier opening questions make the test more accessible and help capture performance at the lower end of the range. More demanding questions can provide better differentiation at the upper end.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Together, this mix allows the test to measure participants across a broad range of performance. The questions are not concentrated at a single difficulty level. Whether a particular question actually distinguishes well between lower and higher performers is examined through its item-total correlation.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Which questions distinguish most clearly between participants?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">A question can have an appropriate level of difficulty and still contribute little to the overall test. I therefore also calculated its <strong>corrected item-total correlation<\/strong>, commonly used as a measure of item discrimination.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This value shows whether people who answer a particular question correctly also tend to score higher on the other 45 questions. \u201cCorrected\u201d means that the question being examined is removed from the comparison score, so it cannot artificially increase its own correlation. [8]<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"576\" src=\"https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-item-total-correlation-43651-1024x576.png\" alt=\"Corrected item-total correlation for each question. of the IQ-Test.\" class=\"wp-image-252\" srcset=\"https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-item-total-correlation-43651-1024x576.png 1024w, https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-item-total-correlation-43651-300x169.png 300w, https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-item-total-correlation-43651-768x432.png 768w, https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-item-total-correlation-43651-1536x864.png 1536w, https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-item-total-correlation-43651.png 1600w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\"><em>Figure 2: Corrected item-total correlation for each question. The line at 0.30 is a rough reference point, not a fixed quality threshold. Based on my analysis. Some information has been obscured to protect the test; further details are available on request.<\/em><\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">For 32 of the 46 questions, the corrected item-total correlation is at least 0.30. The median is <strong>0.362<\/strong>, and q38 has the highest value at <strong>0.562<\/strong>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This means that around 70% of the questions meet the reference value of 0.30 used here. I treat values below 0.20 as a reason for closer review. These are not rigid cutoffs: Very easy or very difficult questions can have lower correlations because there is less variation in the responses, yet they may still serve a useful purpose in the test.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Question 35 (q35) stands out with a value of just <strong>0.017<\/strong>. Whether a participant solved this question was therefore almost unrelated to their performance on the rest of the test. Question 11 (q11) is also comparatively weak at <strong>0.067<\/strong>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Low values can have several causes, including unclear presentation, more than one plausible solution, an unsuitable answer key or a response pattern that differs substantially from the rest of the test. I will therefore review q35 and q11 and then decide whether to retain, revise or replace them.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The overall test does not depend on any single question. Removing q35 would raise alpha only slightly, from 0.876 to <strong>0.879<\/strong>. The finding is therefore mainly a specific opportunity to improve the question pool.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Do the questions measure a common factor?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Reliability alone does not show whether the responses share an underlying structure. To examine this, I conducted an exploratory factor analysis.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Factor analysis looks for patterns: Which questions tend to be solved by the same people, and can these relationships be described by one or more common performance factors? A factor is not a directly observed trait. It is a statistical summary of shared response patterns.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Because each answer is recorded only as correct or incorrect, I based the factor analysis on tetrachoric correlations. This approach assumes that an underlying continuous ability or tendency to solve the problem gives rise to the observed binary response. Tetrachoric correlations are commonly used when conducting factor analyses of binary or ordinal items. [9]<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"576\" src=\"https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-factor-structure-43651-1024x576.png\" alt=\"Eigenvalues of the tetrachoric correlation matrix.\" class=\"wp-image-251\" srcset=\"https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-factor-structure-43651-1024x576.png 1024w, https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-factor-structure-43651-300x169.png 300w, https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-factor-structure-43651-768x432.png 768w, https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-factor-structure-43651-1536x864.png 1536w, https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-factor-structure-43651.png 1600w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\"><em>Figure 3: Eigenvalues of the tetrachoric correlation matrix. The clear gap between the first and subsequent factors indicates a dominant common factor. Based on my analysis.<\/em><\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The first eigenvalue is <strong>12.98<\/strong>, while the second is <strong>2.45<\/strong>. The first is therefore approximately 5.3 times as large as the second. This indicates a common performance factor with a particularly strong influence on the response patterns. The first factor accounts for 28.2% of the total variance in the tetrachoric correlation matrix. This is not a measure of test reliability and does not indicate how reliable the test is overall.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In the one-factor solution, 39 of the 46 questions have a loading of at least 0.30. Participants who perform well in one part of the test therefore also tend to perform better on other questions. This provides a reasonable foundation for calculating one overall score.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Not every question follows exactly the same response pattern, so the analysis does not indicate a perfectly unidimensional test. I did not investigate separate ability domains in this analysis.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Several eigenvalues are greater than 1, but I do not interpret this automatically as evidence of several distinct abilities. The simple rule of retaining every factor with an eigenvalue above 1 can produce too many small factors when a test contains many items. [10]<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Overall, the data support a dominant common performance factor and therefore support the use of a total score. They do not, by themselves, prove that this score is an exact measure of general intelligence.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">What I am taking from the analysis<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">For me, the analysis provides more than a set of summary statistics. It also identifies specific ways to improve the test:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>I will review q35.<\/strong> It is almost unrelated to performance on the rest of the test.<\/li>\n\n\n\n<li><strong>I will also take a closer look at q11.<\/strong> Its corrected item-total correlation is considerably lower than that of most other questions.<\/li>\n\n\n\n<li><strong>I will not evaluate q40 psychometrically again until enough properly stored raw responses are available.<\/strong> The historical correct\/incorrect values are not sufficient for this purpose.<\/li>\n\n\n\n<li><strong>I intend to preserve the broad range of difficulty.<\/strong> It helps the test distinguish between participants at both the lower and upper ends of the performance range.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The current level of quality is not an endpoint. I will continue to monitor, review and improve the test as new data become available.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">What this analysis cannot show<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The large sample allows precise estimates within this dataset, but it does not remove the study&#8217;s methodological limitations.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">No external validation<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">I did not compare these results with scores from an established, professionally administered intelligence test. The analysis therefore does not show how closely the reported IQ score would agree with a score obtained in a clinical or diagnostic setting. If an opportunity for a suitable study arises, I would like to conduct such a comparison in the future.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">No test-retest reliability<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Internal consistency examines how the questions work together within a single test session. This analysis did not examine whether the same person would obtain a similar result when taking the test again under comparable conditions.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">No representative population sample<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Participants chose to take the online test themselves. They do not necessarily represent the general population in terms of age, education, cultural background or other characteristics. This analysis alone therefore cannot establish a new representative IQ norm.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">No statement about the precision of an individual score<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">A reliability coefficient of approximately 0.88 does not mean that every reported IQ score is \u201c88% correct.\u201d Individual interpretation would also require, among other things, sound norming and a standard error of measurement expressed on the IQ scale being used.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">No complete assessment of validity<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Reliability is an important requirement for meaningful measurement, but it is not the same as validity. Validity concerns whether the interpretation and intended use of a test score are supported by appropriate evidence. Reliability, validity and fairness are therefore connected but distinct requirements for a test. [1]<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Conclusion: strong results, with clear limits<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Across 43,651 historical test sessions, the 46 usable questions in the IQ test on IQ-Test.org show high internal consistency. Cronbach&#8217;s alpha (0.876) and McDonald&#8217;s omega (0.875) lead to almost exactly the same conclusion. Most questions are meaningfully related to performance on the rest of the test, cover a broad range of difficulty and contribute to a dominant common performance factor.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">At the same time, the analysis reveals where the test can be improved. I will review q35 in particular, as well as q11 to a lesser extent. Question 40 remains excluded from the historical calculations.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">As an internal overall scale, the test performs reliably and distinguishes across a broad difficulty range. Whether its reported IQ scores agree with professionally administered intelligence tests can only be established through a separate external validation study.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">If you have access to a professionally administered intelligence test and are interested in conducting such a study, I would be pleased to hear from you.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Methodology and calculations in detail<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The following section documents the calculations for readers who want to examine the technical details more closely. It is not required in order to understand the main article.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Sample and data preparation<\/h3>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Data source: historical database export through September 1, 2026<\/li>\n\n\n\n<li>Period covered: May 10, 2010 to September 1, 2026<\/li>\n\n\n\n<li>Inclusion criterion: internally reviewed and valid records only (<code>norm_valid = 1<\/code>)<\/li>\n\n\n\n<li>Sample size: <em>N<\/em> = 43,651<\/li>\n\n\n\n<li>Original number of questions: 47<\/li>\n\n\n\n<li>Questions analyzed: 46<\/li>\n\n\n\n<li>Exclusion: q40, due to a historically inconsistent answer key and unavailable raw responses<\/li>\n\n\n\n<li>Coding: correct = 1, incorrect = 0<\/li>\n\n\n\n<li>Score analyzed: unweighted sum of the 46 item values<\/li>\n<\/ul>\n\n\n\n<h3 class=\"wp-block-heading\">Cronbach&#8217;s alpha<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Alpha was calculated for the binary item values using the classical formula:<\/p>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><code>\u03b1 = [k \/ (k \u2212 1)] \u00d7 [1 \u2212 (\u03a3\u1d62\u208c\u2081\u1d4f \u03c3\u1d62\u00b2 \/ \u03c3\u2093\u00b2)]<\/code><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Here, <em>k<\/em> is the number of questions, <em>\u03c3\u1d62\u00b2<\/em> is the variance of each question and <em>\u03c3\u2093\u00b2<\/em> is the variance of the total score. For binary items, this calculation is equivalent to the KR-20 approach. [2, 3]<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The following table shows a commonly cited rule-of-thumb interpretation:<\/p>\n\n\n\n<figure class=\"wp-block-table is-style-stripes\"><table class=\"has-fixed-layout\"><tbody><tr><th class=\"has-text-align-left\" data-align=\"left\">Alpha<\/th><th>Common label<\/th><\/tr><tr><td class=\"has-text-align-left\" data-align=\"left\">&gt; 0.90<\/td><td>Excellent<\/td><\/tr><tr><td class=\"has-text-align-left\" data-align=\"left\">&gt; 0.80<\/td><td>Good<\/td><\/tr><tr><td class=\"has-text-align-left\" data-align=\"left\">&gt; 0.70<\/td><td>Acceptable<\/td><\/tr><tr><td class=\"has-text-align-left\" data-align=\"left\">&gt; 0.60<\/td><td>Questionable<\/td><\/tr><tr><td class=\"has-text-align-left\" data-align=\"left\">&gt; 0.50<\/td><td>Poor<\/td><\/tr><tr><td class=\"has-text-align-left\" data-align=\"left\">\u2264 0.50<\/td><td>Unacceptable<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Rule-of-thumb labels for alpha, adapted from George and Mallery (2003). [11]<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The result is <strong>\u03b1 = 0.8761<\/strong>. Fixed labels such as \u201cgood\u201d or \u201cexcellent\u201d are only rough guidelines and must be considered in relation to the purpose of the test. Taken together with omega, the corrected item-total correlations and the factor analysis, this value indicates high internal consistency. I do not treat alpha as evidence of unidimensionality or validity. [4]<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">McDonald&#8217;s omega<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">In addition to Cronbach&#8217;s alpha, I calculated McDonald&#8217;s omega. I used a one-factor model based on the observed Pearson correlations among the 46 binary-scored questions. For a standardized one-factor model, the basic principle can be written as:<\/p>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><code>\u03c9 = (\u03a3 \u03bb\u1d62)\u00b2 \/ [(\u03a3 \u03bb\u1d62)\u00b2 + \u03a3 \u03b8\u1d62]<\/code><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Here, <em>\u03bb\u1d62<\/em> is the factor loading of a question, indicating how strongly it is related to the common factor. <em>\u03b8\u1d62<\/em> is the proportion of its variance that is not explained by that factor.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This produces <strong>\u03c9 = 0.8751<\/strong>, almost identical to Cronbach&#8217;s alpha of 0.8761. Both calculations therefore lead to the same overall interpretation. [5, 6]<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Because the questions are recorded only as correct or incorrect, I also conducted an ordinal sensitivity analysis using tetrachoric rather than ordinary Pearson correlations. In simple terms, this model assumes that a continuous ability or tendency to solve the task underlies each observed binary response.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Under these model assumptions, omega is higher at 0.9377. This does not mean that the test suddenly becomes more reliable. The higher value mainly reflects the additional assumptions built into the model. I therefore report it only as a sensitivity analysis, not as a replacement for the observed-data omega of 0.8751. [7]<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Item difficulty<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">For each question <em>i<\/em>, I calculated the proportion of correct answers:<\/p>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><em>p\u1d62 = number of correct answers to question i \/ N<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Here, <em>N<\/em> is the number of test sessions included in the analysis.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Under this definition, a high <em>p\u1d62<\/em> value indicates an easier question. Values range from <strong>0.0219<\/strong> to <strong>0.9813<\/strong>, with a median of <strong>0.5376<\/strong>.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Corrected item-total correlation<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The corrected item-total correlation indicates how well a single question distinguishes between people with higher and lower overall test performance. I calculated it as the Pearson correlation between the response to one question and the total score across the remaining 45 questions:<\/p>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><code>r\u1d62\u209c(i) = cor(X\u1d62, Xtotal \u2212 X\u1d62)<\/code><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The symbols mean:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><em>r\u1d62\u209c(i)<\/em> is the corrected item-total correlation for question <em>i<\/em>.<\/li>\n\n\n\n<li><em>cor<\/em> denotes the Pearson correlation.<\/li>\n\n\n\n<li><em>X\u1d62<\/em> is the response to that question, coded 1 for correct and 0 for incorrect.<\/li>\n\n\n\n<li><em>Xtotal<\/em> is the total score across all 46 questions included in the analysis.<\/li>\n\n\n\n<li><em>Xtotal \u2212 X\u1d62<\/em> is the rest score: the total score without the question currently being examined.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Pearson&#8217;s correlation describes the strength and direction of the relationship between two variables. It ranges from \u22121 to +1. A high positive value here means that participants who answer the question correctly also tend to score higher on the remaining questions. A value close to zero indicates little or no clear relationship. A negative value would be a warning sign because stronger overall performers would then tend to answer the question incorrectly more often.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Because <em>X\u1d62<\/em> can take only the values 0 and 1, this is technically a point-biserial correlation, which is mathematically a special case of Pearson&#8217;s correlation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Removing <em>X\u1d62<\/em> from the total score prevents an artificial increase in the correlation. If the question remained in the total, it would partly be correlated with itself. [8]<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The median corrected item-total correlation is <strong>0.3615<\/strong>. Of the 46 questions, 32 reach at least 0.30 and 11 fall below 0.20. These thresholds are rough reference points, not fixed rules. A low item-total correlation alone is not sufficient reason to remove a question automatically.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Factor analysis: What common structure do the questions show?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">I used factor analysis to examine whether responses to the 46 questions were mainly shaped by a common factor. Factor analysis searches for shared patterns in the responses. A factor is a statistical commonality detected across several questions.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I estimated tetrachoric correlations for all 1,035 pairs of questions. Although each response is observed only as correct or incorrect, the model assumes an underlying continuous ability or tendency to solve the problem that produces a correct response once a certain threshold is crossed. For responses with more than two ordered categories, the related method is known as polychoric correlation. [9]<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The eigenvalues of the tetrachoric correlation matrix provide an initial view of the common structure. They indicate how much of the relationship among the questions can be summarized by each factor. The four largest eigenvalues are:<\/p>\n\n\n\n<ol start=\"1\" class=\"wp-block-list\">\n<li><strong>12.9766<\/strong><\/li>\n\n\n\n<li><strong>2.4518<\/strong><\/li>\n\n\n\n<li><strong>1.9117<\/strong><\/li>\n\n\n\n<li><strong>1.3756<\/strong><\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">The first eigenvalue is more than five times as large as the second and accounts for <strong>28.21%<\/strong> of the total variance. See the main article above for context.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This percentage must not be confused with test reliability. It does not mean that the test is only 28.21% reliable. It describes how much of the overall correlation structure among the 46 questions can be summarized by a single factor.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I also examined how strongly each question was related to this factor. This relationship is called a factor loading. In the one-factor solution, <strong>39 of the 46 questions<\/strong> have a loading of at least 0.30. Question 38 has the highest loading at <strong>0.746<\/strong>, while q35 has a loading of only <strong>0.016<\/strong> and is essentially unrelated to the factor.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Overall, the presence of a common factor is a positive result for the way this test is scored. Because the test produces a total score, its questions should share a common performance basis. The analysis shows that this is true for the large majority of questions.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Several eigenvalues exceed 1, but I do not automatically interpret them as separate abilities. The simple eigenvalue-greater-than-one rule can retain too many small factors, particularly when many questions are analyzed. I therefore do not use it in isolation. [10]<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Stability over time<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">As an additional sensitivity check, I divided the dataset chronologically into two equally sized halves. Alpha was <strong>0.873<\/strong> in the earlier half and <strong>0.880<\/strong> in the later half. Item difficulty values correlated <strong>0.986<\/strong> between the two periods, while the corrected item-total correlations correlated <strong>0.989<\/strong>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The main results are therefore very similar across the two time periods. Individual questions may nevertheless change over time: The percentage of correct answers for q5 fell by approximately 14.4 percentage points between the two halves. In the future, I plan to monitor changes like this using clearly versioned questions and scoring data.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Reproducible results<\/h3>\n\n\n\n<figure class=\"wp-block-table is-style-stripes\"><table class=\"has-fixed-layout\"><tbody><tr><th>Analysis<\/th><th class=\"has-text-align-right\" data-align=\"right\">Value<\/th><\/tr><tr><td>Test sessions<\/td><td class=\"has-text-align-right\" data-align=\"right\">43,651<\/td><\/tr><tr><td>Questions<\/td><td class=\"has-text-align-right\" data-align=\"right\">46<\/td><\/tr><tr><td>Mean total score<\/td><td class=\"has-text-align-right\" data-align=\"right\">24.4468<\/td><\/tr><tr><td>Standard deviation of total score<\/td><td class=\"has-text-align-right\" data-align=\"right\">7.9980<\/td><\/tr><tr><td>Median total score<\/td><td class=\"has-text-align-right\" data-align=\"right\">25<\/td><\/tr><tr><td>Cronbach&#8217;s alpha<\/td><td class=\"has-text-align-right\" data-align=\"right\">0.876134<\/td><\/tr><tr><td>Standardized alpha<\/td><td class=\"has-text-align-right\" data-align=\"right\">0.872617<\/td><\/tr><tr><td>McDonald&#8217;s omega, observed Pearson correlations<\/td><td class=\"has-text-align-right\" data-align=\"right\">0.875077<\/td><\/tr><tr><td>Ordinal omega, tetrachoric sensitivity analysis<\/td><td class=\"has-text-align-right\" data-align=\"right\">0.937743<\/td><\/tr><tr><td>Median item difficulty<\/td><td class=\"has-text-align-right\" data-align=\"right\">0.537605<\/td><\/tr><tr><td>Median corrected item-total correlation<\/td><td class=\"has-text-align-right\" data-align=\"right\">0.361538<\/td><\/tr><tr><td>First tetrachoric eigenvalue<\/td><td class=\"has-text-align-right\" data-align=\"right\">12.9766<\/td><\/tr><tr><td>Second tetrachoric eigenvalue<\/td><td class=\"has-text-align-right\" data-align=\"right\">2.4518<\/td><\/tr><tr><td>Ratio of first to second eigenvalue<\/td><td class=\"has-text-align-right\" data-align=\"right\">5.29<\/td><\/tr><tr><td>Variance accounted for by the first factor<\/td><td class=\"has-text-align-right\" data-align=\"right\">28.21%<\/td><\/tr><tr><td>Loadings on factor 1 \u2265 0.30<\/td><td class=\"has-text-align-right\" data-align=\"right\">39 of 46<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">Download the analysis<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">I am also making the results available in two downloadable formats:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>PDF report (currently available in German):<\/strong> A summary of the dataset, methodology, main findings, charts, limitations and sources.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-analyse-43651-oeffentliche-fassung.pdf\"><strong>Download the PDF report<\/strong><\/a><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Excel workbook (currently using German labels):<\/strong> Key results, anonymized distributions and the eigenvalue spectrum for readers who want to examine the results in more detail.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/iq-test.org\/blog\/wp-content\/uploads\/2026\/09\/iq-test-analyse-43651-oeffentliche-daten-1.xlsx\"><strong>Download the Excel workbook<\/strong><\/a><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">To protect the test, the downloads do not contain individual response data, complete correlation matrices or the raw-score-to-IQ conversion.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Sources and further reading<\/h2>\n\n\n\n<ol start=\"1\" class=\"wp-block-list\">\n<li>American Educational Research Association, American Psychological Association &amp; National Council on Measurement in Education (2014): <em><a href=\"https:\/\/www.aera.net\/publications\/books\/standards-for-educational-psychological-testing-2014-edition\" target=\"_blank\" rel=\"noopener\">Standards for Educational and Psychological Testing<\/a><\/em>.<\/li>\n\n\n\n<li>Cronbach, L. J. (1951): <em><a href=\"https:\/\/doi.org\/10.1007\/BF02310555\" target=\"_blank\" rel=\"noopener\">Coefficient Alpha and the Internal Structure of Tests<\/a><\/em>. <em>Psychometrika, 16<\/em>, 297\u2013334.<\/li>\n\n\n\n<li>Kuder, G. F. &amp; Richardson, M. W. (1937): <em><a href=\"https:\/\/doi.org\/10.1007\/BF02288391\" target=\"_blank\" rel=\"noopener\">The Theory of the Estimation of Test Reliability<\/a><\/em>. <em>Psychometrika, 2<\/em>, 151\u2013160.<\/li>\n\n\n\n<li>Sijtsma, K. (2009): <em><a href=\"https:\/\/pmc.ncbi.nlm.nih.gov\/articles\/PMC2792363\/\" target=\"_blank\" rel=\"noopener\">On the Use, the Misuse, and the Very Limited Usefulness of Cronbach&#8217;s Alpha<\/a><\/em>. <em>Psychometrika, 74<\/em>, 107\u2013120.<\/li>\n\n\n\n<li>McDonald, R. P. (1978): <em><a href=\"https:\/\/doi.org\/10.1177\/001316447803800111\" target=\"_blank\" rel=\"noopener\">Generalizability in Factorable Domains: Domain Validity and Generalizability<\/a><\/em>. <em>Educational and Psychological Measurement, 38<\/em>, 75\u201379.<\/li>\n\n\n\n<li>Dunn, T. J., Baguley, T. &amp; Brunsden, V. (2014): <em><a href=\"https:\/\/doi.org\/10.1111\/bjop.12046\" target=\"_blank\" rel=\"noopener\">From Alpha to Omega: A Practical Solution to the Pervasive Problem of Internal Consistency Estimation<\/a><\/em>. <em>British Journal of Psychology, 105<\/em>, 399\u2013412.<\/li>\n\n\n\n<li>Green, S. B. &amp; Yang, Y. (2009): <em><a href=\"https:\/\/doi.org\/10.1007\/s11336-008-9099-3\" target=\"_blank\" rel=\"noopener\">Reliability of Summed Item Scores Using Structural Equation Modeling: An Alternative to Coefficient Alpha<\/a><\/em>. <em>Psychometrika, 74<\/em>, 155\u2013167.<\/li>\n\n\n\n<li>Henrysson, S. (1963): <em><a href=\"https:\/\/doi.org\/10.1007\/BF02289618\" target=\"_blank\" rel=\"noopener\">Correction of Item-Total Correlations in Item Analysis<\/a><\/em>. <em>Psychometrika, 28<\/em>, 211\u2013218.<\/li>\n\n\n\n<li>Flora, D. B. &amp; Curran, P. J. (2004): <em><a href=\"https:\/\/pmc.ncbi.nlm.nih.gov\/articles\/PMC3153362\/\" target=\"_blank\" rel=\"noopener\">An Empirical Evaluation of Alternative Methods of Estimation for Confirmatory Factor Analysis With Ordinal Data<\/a><\/em>. <em>Psychological Methods, 9<\/em>, 466\u2013491.<\/li>\n\n\n\n<li>Zwick, W. R. &amp; Velicer, W. F. (1986): <em><a href=\"https:\/\/doi.org\/10.1037\/0033-2909.99.3.432\" target=\"_blank\" rel=\"noopener\">Comparison of Five Rules for Determining the Number of Components to Retain<\/a><\/em>. <em>Psychological Bulletin, 99<\/em>, 432\u2013442.<\/li>\n\n\n\n<li>George, D. &amp; Mallery, P. (2003): <em><a href=\"https:\/\/books.google.com\/books?id=AghHAAAAMAAJ\" target=\"_blank\" rel=\"noopener\">SPSS for Windows Step by Step: A Simple Guide and Reference, 11.0 Update<\/a><\/em> (4th ed.). Boston: Allyn &amp; Bacon.<\/li>\n<\/ol>\n","protected":false},"excerpt":{"rendered":"<p>In short: I analyzed 43,651 historical test sessions from IQ-Test.org. The result is strong: Taken together, the 46 questions included in the analysis show remarkably consistent performance. Cronbach&#8217;s alpha and McDonald&#8217;s omega are both approximately 0.88. Most individual questions are also meaningfully related to the overall score, and the response patterns reveal a clear common [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":249,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[25,209],"tags":[197,37,200,203,139],"class_list":["post-246","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-the-science-of-iq","category-iq-test-en","tag-factor-analysis-en","tag-iq-test","tag-iq-test-reliability-en","tag-item-analysis-en","tag-psychometrics"],"_links":{"self":[{"href":"https:\/\/iq-test.org\/blog\/wp-json\/wp\/v2\/posts\/246","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/iq-test.org\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/iq-test.org\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/iq-test.org\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/iq-test.org\/blog\/wp-json\/wp\/v2\/comments?post=246"}],"version-history":[{"count":3,"href":"https:\/\/iq-test.org\/blog\/wp-json\/wp\/v2\/posts\/246\/revisions"}],"predecessor-version":[{"id":253,"href":"https:\/\/iq-test.org\/blog\/wp-json\/wp\/v2\/posts\/246\/revisions\/253"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/iq-test.org\/blog\/wp-json\/wp\/v2\/media\/249"}],"wp:attachment":[{"href":"https:\/\/iq-test.org\/blog\/wp-json\/wp\/v2\/media?parent=246"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/iq-test.org\/blog\/wp-json\/wp\/v2\/categories?post=246"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/iq-test.org\/blog\/wp-json\/wp\/v2\/tags?post=246"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}