This book uses dummy coding for categorical predictors (see Chapter 13). Each category becomes a 0/1 indicator, the intercept is the mean of the reference group, and each coefficient is a difference from that reference group.
That’s the usual approach in sociology. But you’ll see other approaches in published work, especially in psychology. This appendix shows the main alternative, so you’ll recognize it if you run into it.
Let’s look at Southern and non-Southern states in 1977 and compare them on life expectancy.
state.x77_with_names <-as_tibble(state.x77,rownames ="state")states <-bind_cols(state.x77_with_names,region = state.division) |># both are in base R janitor::clean_names() |># lower case and underscoremutate(south =as.integer( # make a south dummystr_detect(as.character(region),"South")))
Warning in FUN(X[[i]], ...): unable to translate '<U+00C4>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00D6>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00DC>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00E4>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00F6>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00FC>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00DF>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00C6>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00E6>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00D8>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00F8>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00C5>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00E5>' to native encoding
states |>group_by(south) |>summarize(mean =mean(life_exp), n =n())
# A tibble: 2 x 3
south mean n
<int> <dbl> <int>
1 0 71.4 34
2 1 69.7 16
D.1 Two approaches to coding a predictor
The coding of south above is called dummy coding or reference coding and is by far the most common in sociology. But sometimes you will see effect coding (also called deviation or contrast coding) that codes the predictor (-1, 1) instead of (0, 1).
ma <-lm(life_exp ~1+ south, data = states)states <- states |>mutate(south_eff =if_else(south ==1, 1, -1))ma_effect <-lm(life_exp ~ south_eff, data = states)
With effect coding, the intercept is the grand mean, or the overall mean you’d expect if both groups were the same size. And \(b_1\) is half the difference between the two groups.
We can check both of those:
group_means <- states |>group_by(south) |>summarize(m =mean(life_exp)) |>pull(m)c(intercept =unname(coef(ma_effect)[1]),unweighted_average =mean(group_means),overall_mean =mean(states$life_exp))
The \(R^2\), fitted values, residuals, and F test are all the same. The only thing that changes is how the parameters are defined, just like when we used relevel() in Chapter 13. Writing the model a different way doesn’t change the model.
D.3 More than two groups
With \(k\) groups, effect coding generalizes to coding each category’s departure from the unweighted grand mean, which is what the mean would be if all groups were of equal size. In R, this is called contr.sum().
states <- states |>mutate(reg4 =factor(state.region))states |>group_by(reg4) |>summarize(mean =mean(life_exp), n =n())
# A tibble: 4 x 3
reg4 mean n
<fct> <dbl> <int>
1 Northeast 71.3 9
2 South 69.7 16
3 North Central 71.8 12
4 West 71.2 13
contrasts(states$reg4) <-contr.sum(4)m4 <-lm(life_exp ~ reg4, data = states)round(coef(m4), 3)
The intercept is the unweighted grand mean, and each of the three coefficients is how far one region is from it:
region_means <- states |>group_by(reg4) |>summarize(m =mean(life_exp)) |>pull(m)c(intercept =unname(coef(m4)[1]), unweighted_grand_mean =mean(region_means))
intercept unweighted_grand_mean
70.99299 70.99299
round(region_means -mean(region_means), 3)
[1] 0.271 -1.287 0.774 0.242
There are three coefficients for four groups, as usual. The fourth one isn’t shown because we don’t need it. The deviations have to add up to zero, so the last one is just minus the sum of the others.
printed1 printed2 printed3 implied_fourth total
0.2714503 -1.2867441 0.7736725 0.2416213 0.0000000
You can see these estimates sum to zero. That’s why it’s called deviation coding: every group is compared to the same center rather than to one reference group.
NoteWhen you will meet it
Effect coding is common in experimental psychology, where the groups are often the same size and the grand mean makes sense as a baseline. When the groups are the same size, the unweighted and weighted means are the same, so the intercept is just “the mean.”
In sociology, the groups are usually different sizes, and there’s often an obvious reference category (no college, unmarried, non-South, etc.). Dummy coding lets you compare to that group, in units a reader can check against a table of group means.
That’s why this book uses dummy coding. But now if you see a table where the intercept doesn’t match any group’s mean and the coefficients look half as big as you’d expect, you’ll know what’s going on!
D.4 A caution about unbalanced groups
The main thing to watch out for is the grand mean.
With 16 southern and 34 other states, the effect-coded intercept 70.568 is not the mean life expectancy of American states. That’s 70.879. The intercept is the average of the two group means, which gives each state in the smaller group more weight than each state in the bigger group. So if you use effect coding, be sure to say that the intercept is the unweighted grand mean, not the sample mean.