We’re going to look at models with multiple categorical predictors, including their interactions. I’m going to use dummy coding and model comparison rather than contrast coding.
We are going to use the full GSS data for this example. That’s the cumulative file you saved in the setup chapter. We need it here because we want to look at about fifty years of surveys.
d <-readRDS(here::here("data", "gss-1972-2024.rds")) |> haven::zap_labels() |>select(tvhours, degree, year, sex) |>drop_na() |>mutate(female =if_else(sex ==2, 1, 0), # femaledegree_fac =factor(degree), # degree as factoryear_fac =factor(year)) # svy year as factorc(respondents =nrow(d), survey_years =nlevels(d$year_fac))
respondents survey_years
46000 29
14.1 Multiple categorical predictors
We’re going to consider a few categorical predictors of tvhours:
female: whether the respondent is female (0, 1)
degree_fac: highest degree earned from none (0) to graduate degree (4); stored as a factor
year_fac: a factor encoding the survey year1 from 1975 to 2024 (with gaps)
1year is clearly a “numeric” variable in the obvious sense. We have year_fac stored as a factor so that we make no assumptions about its functional (e.g., straight line) relationship to the outcome when we use it in a model. This is the same trade-off we saw in Chapter 13. Here we can afford the extra parameters because we have tens of thousands of cases.
Let’s consider a simple model that uses all of these additively.
the degree_fac ones that show how each level of degree_fac are different from the “none” reference category
the female one that shows how female respondents are different from males
the year_fac ones that show how different each survey is from the 1975 reference
Each one is a difference from a reference category, just like in Chapter 13. There are just a lot of them! With this many coefficients, the table isn’t a very useful way to understand the model.
This is easier to see in picture form. I will use plot_predictions() from the marginaleffects package. Using the newdata = "balanced" argument means that the predictions are averaged over equal values of the other predictors (i.e, in the first plot: half male, equal representation from all survey years).
Figure 14.6: Predicted television hours by sex, balanced over education and survey year.
NoteBalanced predictions are a choice
Averaging over an equal mix of survey years treats every year as equally important, even though the GSS didn’t interview the same number of people every year. And averaging over half men and half women is also something we made up. Balanced predictions describe a hypothetical population that you’ve specified. That’s often what you want for comparisons (because it holds the composition constant), but you should say that’s what you did when you report the results.
14.2 Interactions
The model above says that there are educational differences and year differences and sex differences. But the model also assumes that those differences are constant. That is, m1 assumes that, for example, the educational differences don’t differ by sex or year.
We can relax that assumption in several ways and compare the results. We can allow any pair of those differences to moderate each other (3 options) or all three to affect each other.
2 The AIC_wt and BIC_wt columns (AIC and BIC “weights”) turn the differences into something like relative probabilities. To get them, you take the difference between each model and the best one, multiply by \(-\frac{1}{2}\), exponentiate, and then divide by the total. The BIC weights are the same idea as the posterior probabilities in Chapter 10.
model parameters PRE AIC AIC_wt BIC BIC_wt
1 additive 34 0.0547 214593.7 0 214899.5 0.997
2 degree x sex 38 0.0553 214570.3 1 214911.0 0.003
3 degree x year 146 0.0580 214657.5 0 215941.7 0.000
4 sex x year 62 0.0556 214602.9 0 215153.3 0.000
5 all three 290 0.0616 214766.6 0 217308.9 0.000
PRE goes up with every interaction we add (as always), but the three-way model with 290 parameters is the worst according to both AIC and BIC.
With 46,000 cases, it’s not too surprising that the BIC prefers the simple additive model. The AIC, however, prefers the model where degree and sex moderate each other. Let’s take a look at that.
This disagreement isn’t a problem. As we saw in Chapter 10, AIC and BIC are trying to answer different questions, and BIC’s penalty gets bigger with the sample size. Here \(\log(n)\) is about 10.7, so BIC’s penalty per parameter is more than three times as big as AIC’s. An interaction can help with prediction without being something we’d want to claim is “real.”
# A tibble: 5 x 4
degree_fac f0 f1 gap
<fct> <dbl> <dbl> <dbl>
1 0 3.63 4.05 0.424
2 1 3.04 3.22 0.185
3 2 2.64 2.71 0.076
4 3 2.24 2.26 0.02
5 4 1.91 1.96 0.049
Among respondents with no degree, women are predicted to watch about 0.424 hours more television per day than men. Among those with graduate degrees, the gap is about 0.049 hours (basically nothing). So the gap between men and women gets smaller as education goes up. The additive model can’t show that, since it assumes the sex gap is the same at every level of education.
WarningSay what the model says
It is tempting to write that “education reduces the sex gap in television watching.” But that’s not what we found. We compared people with different amounts of education in a survey. We didn’t change anyone’s education, and lots of other things (birth cohort, employment, household composition, etc.) are different across these groups too.
A better way to say it: among respondents with more education, the predicted difference between women and men is smaller.
14.3 Recap
several categorical predictors each contribute a block of dummy variables measured against their own reference category
beyond a handful of coefficients, read a model through predictions rather than a coefficient table
balanced predictions average over the other predictors with equal weight, so they describe a population you have constructed
interactions between categorical predictors let group differences differ across groups
information criteria avoid testing each interaction separately, and they can disagree
a preferred interaction is best reported as predicted values in each cell