Monarchy & Republic in the Laboratory of History by N. Fakhr - HTML preview
Download the book in PDF, ePub for a complete version.
Part Two
Appendix I: Research Methodology and Statistical Analysis
1. General Research Design
The purpose of this study is to compare the performance of countries with monarchical and republican systems of government across a set of measurable indicators of governance, development, and welfare. The unit of analysis is the country, and comparisons are based, wherever possible, on data from the same year and from reputable international sources.
This study is observational and comparative in nature. Countries have not been randomly assigned to different political systems, and numerous historical, geographical, cultural, economic, and institutional factors may be associated both with the form of government and with the indicators examined. The results should therefore be understood primarily as observed empirical differences between groups of countries and should not, without additional evidence, be interpreted as proof of a causal relationship between the form of government and the outcomes under study.
The central principle of this methodology is to avoid beginning with the assumption that either form of government is inherently superior to the other. Instead, the actual performance of countries is examined using objective indicators, after which the magnitude of the differences between the groups is assessed.
2. Unit of Analysis and Classification of Countries
Countries are classified into two main groups according to their political structure:
- Monarchies: countries in which the head of state is a king, queen, emir, sultan, or holder of a comparable hereditary title.
- Republics: countries in which the office of head of state is not hereditary and the formal structure of government is republican.
For more detailed analysis, monarchies are divided in some comparisons into two subgroups: constitutional monarchies and non-constitutional monarchies. In this book, the latter category includes monarchies in which the monarch exercises greater political authority than in parliamentary constitutional monarchies, including absolute monarchies and certain semi-constitutional systems.
Commonwealth realms are also examined separately. These are independent states that recognize the British monarch as their head of state. Because of their distinctive institutional structure and shared monarch, they are treated as a separate group in the analyses rather than being combined with independent monarchies.
This distinction also prevents overlap between groups in the statistical comparisons.
3. Geographical Coverage
Comparisons are first conducted at the global level and then, where the number of countries permits meaningful analysis, at the continental level.
Regional analyses cover Asia, Europe, Africa, the Americas, and Oceania. The Middle East is also examined separately. It is treated as a distinct analytical region because its countries share certain relative similarities in history, geography, climate, culture, religion, and, for some indicators, economic structure and natural resources. From this perspective, the Middle East serves to some extent as a “natural laboratory” for comparing political systems, although this expression should not be understood to imply the controlled conditions of a true experiment.
Where the number of countries in a group is very small, inferential interpretation is avoided or the findings are reported with greater caution. For example, when a group contains only one country, within-group dispersion cannot be calculated and conventional tests of differences between means cannot be validly performed.
4. Data Sources and Missing Data
For each indicator, data are drawn from reputable international sources, and the source and year are specified in the corresponding section of the book. For indicators such as GDP per capita and GDP per capita based on purchasing power parity, World Bank data form the basis of the calculations.
The general principle of the study is to use the latest common and comparable dataset available for the countries under examination in the selected year.
If the primary source contains no data for a particular country, its value is neither estimated nor imputed from similar countries. Instead, that country is excluded from the calculation for that particular indicator. Consequently, the number of countries included in the analysis may vary from one indicator to another.
This point is especially important when interpreting the results, because missing data are not necessarily random, and the exclusion of particular countries may affect group means and other statistical measures.
5. Descriptive Statistics
For each group, at least three principal descriptive statistics are calculated:
- Mean
- Median
- Sample standard deviation
The mean provides a measure of the general level of an indicator, but in cross-country data it may be strongly affected by a small number of exceptionally high or low values. For this reason, the median is also reported alongside the mean. Considering the two measures together helps determine whether an observed difference is driven mainly by a few exceptional countries or is also evident near the center of the distribution.
The sample standard deviation is calculated to show the degree of dispersion of countries around the group mean. Where comparison of relative dispersion is important, the coefficient of variation is also used.
6. Tests of Differences Between Means
Differences between the means of two independent groups are tested using Welch’s t-test.
Welch’s test was chosen because, unlike the conventional Student’s t-test under the assumption of equal variances, it does not require the two groups to have equal variances and performs well when group sizes differ. This is particularly important in the present study because republics generally far outnumber monarchies, and the dispersion of the indicators is not necessarily equal across the two groups.
The null hypothesis is that the means of the two groups are equal.
Where the research hypothesis specifies a direction in advance—for example, testing the hypothesis that the mean value of a particular indicator is higher among monarchies than among republics—a one-tailed test is reported. Where the question concerns simply whether the two groups differ, without a direction specified in advance, a two-tailed test is appropriate. The direction of the test must be determined before the result is interpreted and must not be changed in response to the observed outcome.
A conventional significance level of 5 percent (α = 0.05) is used as the primary criterion for statistical significance. However, exact p-values are also reported so that the results are not reduced merely to a binary distinction between “significant” and “not significant.”
7. Effect Size
Statistical significance alone does not indicate the magnitude or practical importance of a difference. A difference may be statistically significant but practically small; conversely, a relatively large difference may be observed in a sample but fail to reach the conventional threshold of statistical significance because the number of countries is small.
For this reason, Hedges’ g is calculated alongside Welch’s test.
Hedges’ g standardizes the difference between two group means relative to the dispersion of the data and, particularly in small samples, provides a more appropriate correction for sampling bias than Cohen’s d.
In interpreting effect sizes, values close to zero indicate negligible differences, while larger absolute values of g indicate greater differences. These thresholds are treated as approximate interpretive guidelines rather than as fixed or natural boundaries.
The sign of the effect size indicates the direction of the difference. In the book’s main comparisons, a positive sign means that the mean for monarchies is higher than the mean for republics.
8. Analysis of Data Distributions
Because comparisons of means alone cannot reveal the shape of the distribution of countries, Kernel Density Estimation (KDE) is used for appropriate indicators.
Unlike a histogram, a KDE plot does not depend on an arbitrary division of observations into bins. Instead, it provides a continuous estimate of the approximate distribution of the data.
For economic variables such as GDP per capita and GDP per capita based on purchasing power parity, whose distributions are strongly right-skewed and in which values for wealthy countries may be many times those of lower-income countries, KDE is calculated using the base-10 logarithm of the data. The horizontal axis is then converted back to the actual values of the indicator so that the graph remains easy for readers to interpret.
The same standardized rule for bandwidth selection is applied to all groups shown in a given chart. In the charts presented in this book, Scott’s rule is used to determine bandwidth. The shapes of the curves are therefore direct products of the underlying data and a consistent computational method rather than curves drawn or adjusted manually.
9. Regional Analyses and Robustness Checks
Global comparisons may be influenced by historical, geographical, and economic differences across regions. For this reason, the analyses are repeated at the continental level and for the Middle East. The purpose of these comparisons is to reduce—not eliminate—environmental heterogeneity and to examine whether the global pattern is also observed across different regions.
For Europe, a robustness check is also conducted in which republics with a history of communist rule are excluded from the European republican group, and monarchies are compared with European republics that have never experienced communist rule.
This analysis should not be interpreted as implying that excluding countries with a communist past necessarily brings us closer to identifying the “pure effect” of republican government. A country’s historical path into communism may itself have been associated with its political structure and institutional development. The robustness check therefore addresses a more limited question: if republics with a history of communist rule are removed from the comparison, how much does the observed difference between monarchies and the remaining republics change?
Accordingly, the robustness analysis does not replace the primary comparison; it complements it.
10. Limitations of Interpretation
The findings of this study should be interpreted in light of several important limitations.
First, the study is observational. Therefore, the existence of differences between monarchies and republics does not by itself establish a causal relationship between the form of government and the outcome being examined.
Second, the numbers of countries in the two groups are unequal, and in some regions the number of monarchies or Commonwealth realms is very small. This increases uncertainty in the estimates and reduces the statistical power of the tests.
Third, countries are not units that are entirely independent of one another historically or geographically. Geographic proximity, colonial history, wars, trade, natural resources, regional institutions, and the diffusion of political and economic models can create interdependence among countries, whereas conventional statistical tests generally assume independence of observations.
Fourth, any classification of countries necessarily simplifies some of the complexity of real political systems. Two countries may both be classified as “republics” or “monarchies” while differing substantially in institutional quality, degree of democracy, concentration of political power, and political history.
Fifth, missing data may affect the results. For example, if data for a particular indicator are unavailable for some exceptionally wealthy or poor countries, the estimated group mean may be lower or higher than it otherwise would have been. For this reason, important instances of missing data are explicitly reported in the text.
11. Governing Principle for Interpreting the Results
This study distinguishes among three levels of evidence: descriptive differences, statistical significance, and effect size.
A higher mean or median in one group does not by itself establish the existence of a persistent difference in the broader population of countries. Conversely, a p-value that does not cross the conventional threshold of 0.05 does not mean that there is “no difference,” particularly when one of the groups contains only a small number of countries.
For this reason, final conclusions are based on the totality of the evidence: the direction and magnitude of differences between means, differences between medians, dispersion of the data, p-values, effect sizes, the shapes of the distributions, and the consistency of the observed pattern across regional comparisons.
The purpose of this approach is to allow the data to speak before theoretical judgments are made, while at the same time avoiding conclusions that go beyond what the research design can support.
