Appendix D — Effect coding

This book uses dummy coding for categorical predictors (see Chapter 13). Each category becomes a 0/1 indicator, the intercept is the mean of the reference group, and each coefficient is a difference from that reference group.

That’s the usual approach in sociology. But you’ll see other approaches in published work, especially in psychology. This appendix shows the main alternative, so you’ll recognize it if you run into it.

Let’s look at Southern and non-Southern states in 1977 and compare them on life expectancy.

state.x77_with_names <- as_tibble(state.x77,
                                  rownames = "state")

states <- bind_cols(state.x77_with_names,
                    region = state.division) |> # both are in base R
  janitor::clean_names() |>                     # lower case and underscore
  mutate(south = as.integer(                    # make a south dummy
    str_detect(as.character(region),
               "South")))
Warning in FUN(X[[i]], ...): unable to translate '<U+00C4>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00D6>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00DC>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00E4>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00F6>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00FC>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00DF>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00C6>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00E6>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00D8>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00F8>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00C5>' to native encoding
Warning in FUN(X[[i]], ...): unable to translate '<U+00E5>' to native encoding
states |>
  group_by(south) |>
  summarize(mean = mean(life_exp), n = n())
# A tibble: 2 x 3
  south  mean     n
  <int> <dbl> <int>
1     0  71.4    34
2     1  69.7    16

D.1 Two approaches to coding a predictor

The coding of south above is called dummy coding or reference coding and is by far the most common in sociology. But sometimes you will see effect coding (also called deviation or contrast coding) that codes the predictor (-1, 1) instead of (0, 1).

ma <- lm(life_exp ~ 1 + south, data = states)

states <- states |>
  mutate(south_eff = if_else(south == 1, 1, -1))

ma_effect <- lm(life_exp ~ south_eff, data = states)
Show code
msummary(list("MA (dummy)"  = ma,
              "MA (effect)" = ma_effect),
         fmt = 3,
         estimate = "{estimate} ({std.error})",
         statistic = NULL,
         gof_map = c("nobs", "F", "r.squared"))
MA (dummy) MA (effect)
(Intercept) 71.430 (0.185) 70.568 (0.164)
south -1.724 (0.327)
south_eff -0.862 (0.164)
Num.Obs. 50 50
F 27.739 27.739
R2 0.366 0.366

With effect coding, the intercept is the grand mean, or the overall mean you’d expect if both groups were the same size. And \(b_1\) is half the difference between the two groups.

We can check both of those:

group_means <- states |>
  group_by(south) |>
  summarize(m = mean(life_exp)) |>
  pull(m)

c(intercept          = unname(coef(ma_effect)[1]),
  unweighted_average = mean(group_means),
  overall_mean       = mean(states$life_exp))
         intercept unweighted_average       overall_mean 
          70.56827           70.56827           70.87860 
c(effect_coefficient  = unname(coef(ma_effect)[2]),
  half_the_difference = unname(diff(group_means) / 2))
 effect_coefficient half_the_difference 
         -0.8620221          -0.8620221 

The slope is half the difference because going from -1 to +1 is two units, so each unit is half the gap.

D.2 The model is the same model

Look at the nobs, F, and R2 rows of that table. They’re exactly the same!

all.equal(unname(fitted(ma)), unname(fitted(ma_effect)))
[1] TRUE

The \(R^2\), fitted values, residuals, and F test are all the same. The only thing that changes is how the parameters are defined, just like when we used relevel() in Chapter 13. Writing the model a different way doesn’t change the model.

D.3 More than two groups

With \(k\) groups, effect coding generalizes to coding each category’s departure from the unweighted grand mean, which is what the mean would be if all groups were of equal size. In R, this is called contr.sum().

states <- states |>
  mutate(reg4 = factor(state.region))

states |>
  group_by(reg4) |>
  summarize(mean = mean(life_exp), n = n())
# A tibble: 4 x 3
  reg4           mean     n
  <fct>         <dbl> <int>
1 Northeast      71.3     9
2 South          69.7    16
3 North Central  71.8    12
4 West           71.2    13
contrasts(states$reg4) <- contr.sum(4)

m4 <- lm(life_exp ~ reg4, data = states)

round(coef(m4), 3)
(Intercept)       reg41       reg42       reg43 
     70.993       0.271      -1.287       0.774 

The intercept is the unweighted grand mean, and each of the three coefficients is how far one region is from it:

region_means <- states |>
  group_by(reg4) |>
  summarize(m = mean(life_exp)) |>
  pull(m)

c(intercept = unname(coef(m4)[1]), unweighted_grand_mean = mean(region_means))
            intercept unweighted_grand_mean 
             70.99299              70.99299 
round(region_means - mean(region_means), 3)
[1]  0.271 -1.287  0.774  0.242

There are three coefficients for four groups, as usual. The fourth one isn’t shown because we don’t need it. The deviations have to add up to zero, so the last one is just minus the sum of the others.

deviations <- coef(m4)[-1]

c(printed = unname(deviations),
  implied_fourth = unname(-sum(deviations)),
  total = 0)
      printed1       printed2       printed3 implied_fourth          total 
     0.2714503     -1.2867441      0.7736725      0.2416213      0.0000000 

You can see these estimates sum to zero. That’s why it’s called deviation coding: every group is compared to the same center rather than to one reference group.

NoteWhen you will meet it

Effect coding is common in experimental psychology, where the groups are often the same size and the grand mean makes sense as a baseline. When the groups are the same size, the unweighted and weighted means are the same, so the intercept is just “the mean.”

In sociology, the groups are usually different sizes, and there’s often an obvious reference category (no college, unmarried, non-South, etc.). Dummy coding lets you compare to that group, in units a reader can check against a table of group means.

That’s why this book uses dummy coding. But now if you see a table where the intercept doesn’t match any group’s mean and the coefficients look half as big as you’d expect, you’ll know what’s going on!

D.4 A caution about unbalanced groups

The main thing to watch out for is the grand mean.

c(n_south = sum(states$south), n_other = sum(!states$south))
n_south n_other 
     16      34 

With 16 southern and 34 other states, the effect-coded intercept 70.568 is not the mean life expectancy of American states. That’s 70.879. The intercept is the average of the two group means, which gives each state in the smaller group more weight than each state in the bigger group. So if you use effect coding, be sure to say that the intercept is the unweighted grand mean, not the sample mean.