Ordered Logit

Ordered Outcomes

Suppose an ordered outcome \(y_i \in \{1, 2, \ldots, J\}\), where \(\{1, 2, \ldots, J\}\) correspond to qualitative labels that have a natural ordering.

Example 1

The ANES asks respondents to place themselves along a 7-point ideology scale with positions labeled “extremely liberal,” “liberal,” “slightly liberal,” “moderate/middle of the road,” “slightly conservative,” “conservative,” and “extremely conservative.”

We typically let 1, 2, 3, 4, 5, 6, and 7 correspond to these seven ordered labels.

Example 2

The Political Terror Scale (PTS) is a country-year measure of state repression. This measure assigns one of five qualitative labels summarizing the amount of state repression to each country-years based on reports from Amnesty International and the U.S. State Department.

Numeric Coding Qualitative Label
1 Countries under a secure rule of law, people are not imprisoned for their views, and torture is rare or exceptional. Political murders are extremely rare.
2 There is a limited amount of imprisonment for nonviolent political activity. However, few persons are affected, torture and beatings are exceptional. Political murder is rare.
3 There is extensive political imprisonment, or a recent history of such imprisonment. Execution or other political murders and brutality may be common. Unlimited detention, with or without a trial, for political views is accepted.
4 Civil and political rights violations have expanded to large numbers of the population. Murders, disappearances, and torture are a common part of life. In spite of its generality, on this level terror affects those who interest themselves in politics or ideas.
5 Terror has expanded to the whole population. The leaders of these societies place no limits on the means or thoroughness with which they pursue personal or ideological goals.

Ordered Logit

  • Let linear predictor \(\eta_i\) be our usual linear predictor so that \(\eta_i = X_i \beta\), except without an intercept \(\beta_0\).
  • Now define \(J-1\) “cutpoints” \(\alpha_1 < \alpha_2 < \cdots < \alpha_{J-1}\) that act as a separate intercept for each category.
  • The ordered logit defines the the cumulative probabilities as

\[ \Pr(Y_i \leq j) = \text{logit}^{-1}(\alpha_j - \eta_i), \qquad j = 1, \ldots, J-1. \]

Then probabilities for each category are

\[ \begin{align} \Pr(Y_i = 1) &= \text{logit}^{-1}(\alpha_1 - \eta_i) \\ \Pr(Y_i = 2) &= \text{logit}^{-1}(\alpha_2 - \eta_i) - \text{logit}^{-1}(\alpha_{1} - \eta_i) \\ \Pr(Y_i = 3) &= \text{logit}^{-1}(\alpha_3 - \eta_i) - \text{logit}^{-1}(\alpha_{2} - \eta_i) \\ &\vdots\\ \Pr(Y_i = j) &= \overbrace{\text{logit}^{-1}(\alpha_j - \eta_i)}^{\Pr(Y_i \leq j)} - \overbrace{\text{logit}^{-1}(\alpha_{j-1} - \eta_i)}^{\Pr(Y_i \leq j-1)}, \qquad j = 2, \ldots, J-1 \\ &\vdots\\ \Pr(Y_i = J) &= 1 - \text{logit}^{-1}(\alpha_{J-1} - \eta_i). \end{align} \]

Illustrative Figure

The figure on the next slide shows this graphically for an outcome with four possible values.

  • I set the three cutpoints to \(\alpha = (-2.5,-1.3,1.3)\).
  • I set the linear predictor for an example observation to \(\eta_i = -1.5\).

Example: Approval of President Bush, 1992

The data

  • Data: 1992 ANES Time Series Study (Miller et al. 2016).
  • Outcome: approval of President Bush’s job performance, four categories: strongly disapprove, disapprove, approve, and strongly approve.
  • Predictors:
    • military_force: how willing the United States should be to use military force to solve international problems, from 1 (extremely willing) to 5 (never willing).
    • ideology_distance: distance between the respondent’s ideology and where they place Bush, from 0 (same position) to 6 (opposite ends of the 7-point scale).
    • economy: evaluation of the national economy compared with a year ago, from 1 (much better) to 5 (much worse).
    • party_id: party identification, from −3 (strong Democrat) through 0 (independent) to 3 (strong Republican).
    • education: years of education.
  • 750 respondents; 558 complete on these variables.

bush-approval-1992.csv

First, load the data set with the approval ratings (bush_approval) and the five predictors.

# load data
approval <- read_csv("https://pos5747.github.io/data/bush-approval-1992.csv") |>
  glimpse()
Rows: 750
Columns: 10
$ bush_approval     <chr> "4. Strongly approve", "4. Strongly approve", "3. Ap…
$ military_force    <dbl> 3, 2, 3, 1, 3, 3, 3, 2, 3, 1, 4, 4, 2, 2, 3, 1, 3, 2…
$ ideology_distance <dbl> 3, 0, 1, 1, NA, NA, NA, 1, 1, 1, 2, NA, 0, 1, NA, 1,…
$ economy           <dbl> 4, 4, 3, 3, 5, 5, 4, 2, 4, 3, 5, 5, 4, 4, 4, 5, 5, 4…
$ party_id          <dbl> 2, 3, 3, 3, -2, -3, 2, 3, 2, 1, -3, -2, -2, 3, -3, 2…
$ education         <dbl> 17, 8, 16, 11, 13, 12, 16, 17, 13, 17, 11, 8, 12, 16…
$ ideology          <dbl> 3, 6, 6, 6, NA, 4, 6, 7, 6, 7, 2, 5, 3, 5, NA, 5, NA…
$ gulf_war          <dbl> 1, 1, 1, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 1, 1, 1, 0, 1…
$ nonwhite          <dbl> 0, 0, 0, 1, 0, 1, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0…
$ vote_1992         <chr> "Bush", "Bush", "Bush", "Clinton", "Clinton", "Clint…

Notice that the ordinal approval labels have a numerical prefix as part of the character string. This is helpful to (1) quickly remind us of the order and (2) ensure that the alphabetic order and ordinal rankings of their values are the same.

Because the dataset is in .csv format, bush_approval is loaded as a character string, but we can easily convert it to an ordered factor.

# convert the outcome to an ordered factor
approval <- approval |>
  mutate(bush_approval = factor(bush_approval, ordered = TRUE)) |>
  glimpse()
Rows: 750
Columns: 10
$ bush_approval     <ord> 4. Strongly approve, 4. Strongly approve, 3. Approve…
$ military_force    <dbl> 3, 2, 3, 1, 3, 3, 3, 2, 3, 1, 4, 4, 2, 2, 3, 1, 3, 2…
$ ideology_distance <dbl> 3, 0, 1, 1, NA, NA, NA, 1, 1, 1, 2, NA, 0, 1, NA, 1,…
$ economy           <dbl> 4, 4, 3, 3, 5, 5, 4, 2, 4, 3, 5, 5, 4, 4, 4, 5, 5, 4…
$ party_id          <dbl> 2, 3, 3, 3, -2, -3, 2, 3, 2, 1, -3, -2, -2, 3, -3, 2…
$ education         <dbl> 17, 8, 16, 11, 13, 12, 16, 17, 13, 17, 11, 8, 12, 16…
$ ideology          <dbl> 3, 6, 6, 6, NA, 4, 6, 7, 6, 7, 2, 5, 3, 5, NA, 5, NA…
$ gulf_war          <dbl> 1, 1, 1, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 1, 1, 1, 0, 1…
$ nonwhite          <dbl> 0, 0, 0, 1, 0, 1, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0…
$ vote_1992         <chr> "Bush", "Bush", "Bush", "Clinton", "Clinton", "Clint…
# examine levels of ordinal variable
levels(approval$bush_approval)
[1] "1. Strongly disapprove" "2. Disapprove"          "3. Approve"            
[4] "4. Strongly approve"   

Because of the numbers in the prefix, the default alphabetical ordering is exactly the ordering we want.

# fit ordered logit model
library(MASS)  # be wary of conflicts with filter() and select()
fit <- polr(bush_approval ~ military_force + ideology_distance + economy + party_id + education, 
            data = approval, Hess = TRUE)

# summarize fit
summary(fit)
Call:
polr(formula = bush_approval ~ military_force + ideology_distance + 
    economy + party_id + education, data = approval, Hess = TRUE)

Coefficients:
                     Value Std. Error t value
military_force    -0.56048    0.10861  -5.160
ideology_distance -0.36441    0.06611  -5.512
economy           -0.59655    0.09957  -5.991
party_id           0.43999    0.05214   8.438
education         -0.07244    0.03710  -1.953

Intercepts:
                                     Value   Std. Error t value
1. Strongly disapprove|2. Disapprove -6.9981  0.7509    -9.3198
2. Disapprove|3. Approve             -5.5145  0.7286    -7.5684
3. Approve|4. Strongly approve       -3.2725  0.7031    -4.6544

Residual Deviance: 1175.889 
AIC: 1191.889 
(192 observations deleted due to missingness)

Interpreting the coefficients

  • Positive party_id coefficient means that higher categories are more likely, so Republicans are more approving.
  • Negative coefficients indicate who is more disapproving:
    • respondents who are more opposed to military force
    • respondents who are ideologically further from Bush
    • respondents who perceive the economy as worse.

Probabilities with {marginaleffects}

To interpret the coefficients, we can compute the probability of falling into each category as military_force varies from 1 to 5, for strong Democrats (party_id = -3) and strong Republicans (party_id = 3). The other predictors sit at their means (the datagrid() default).

# use {marginaleffects}
library(marginaleffects)
p <- predictions(fit, 
            newdata = datagrid(military_force = 1:5, 
                               party_id = c(-3, 3))) |>
  glimpse()
Rows: 40
Columns: 15
$ rowid             <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 1, 2, 3, 4, 5, 6, 7, …
$ group             <fct> 1. Strongly disapprove, 1. Strongly disapprove, 1. S…
$ estimate          <dbl> 0.27122253, 0.02587275, 0.39461553, 0.04445201, 0.53…
$ std.error         <dbl> 0.058140543, 0.007652858, 0.052186323, 0.010035746, …
$ statistic         <dbl> 4.664947, 3.380796, 7.561666, 4.429368, 11.877471, 5…
$ p.value           <dbl> 3.086971e-06, 7.227631e-04, 3.979397e-14, 9.450975e-…
$ s.value           <dbl> 18.30538, 10.43419, 44.51444, 16.69111, 105.66956, 2…
$ conf.low          <dbl> 0.15726916, 0.01087342, 0.29233222, 0.02478231, 0.44…
$ conf.high         <dbl> 0.38517590, 0.04087207, 0.49689884, 0.06412171, 0.62…
$ military_force    <int> 1, 1, 2, 2, 3, 3, 4, 4, 5, 5, 1, 1, 2, 2, 3, 3, 4, 4…
$ party_id          <dbl> -3, 3, -3, 3, -3, 3, -3, 3, -3, 3, -3, 3, -3, 3, -3,…
$ df                <dbl> Inf, Inf, Inf, Inf, Inf, Inf, Inf, Inf, Inf, Inf, In…
$ economy           <int> 4, 4, 4, 4, 4, 4, 4, 4, 4, 4, 4, 4, 4, 4, 4, 4, 4, 4…
$ education         <int> 14, 14, 14, 14, 14, 14, 14, 14, 14, 14, 14, 14, 14, …
$ ideology_distance <int> 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2…
ggplot(p, aes(x = military_force, y = estimate, color = group)) + 
  geom_point() + 
  geom_line() +
  facet_wrap(vars(party_id), labeller = label_both)

Computing cumulative probabilities

  • However, the probabilities above are somewhat awkward to interpret.
  • As some categories shrink, it isn’t immediately clear what categories are growing as a result. Instead, we can compute an easier-to-interpret cumulative probability \(\Pr(y_i \leq j)\) using cumsum() within arrange()ed groups.
cumulative_p <- p %>%
  arrange(party_id, military_force, group) %>%
  group_by(party_id, military_force) %>%
  mutate(cumulative_prob = cumsum(estimate),
         j = as.integer(group),
         cdf_label = paste0("Pr(Y ≤ ", group, ")"),
         cdf_label = reorder(cdf_label, j)) %>%
  ungroup()
ggplot(cumulative_p, aes(x = military_force, y = cumulative_prob, color = cdf_label)) + 
  geom_line() +
  facet_wrap(vars(party_id), labeller = label_both)

Alternatively, position_stack(reverse = TRUE) with geom_line() or geom_area() makes ggplot() stack the probabilities for us: the line for category \(j\) sits at \(\Pr(Y_i \leq j)\), the same curves as above without computing them by hand. (The default, position = "stack", stacks the first level on top, which gives \(\Pr(Y_i \geq j)\) instead.)

ggplot(p, aes(x = military_force, y = estimate, color = group)) +
  geom_line(position = position_stack(reverse = TRUE)) +
  facet_wrap(vars(party_id), labeller = label_both)
ggplot(p, aes(x = military_force, y = estimate, fill = group)) +
  geom_area(position = position_stack(reverse = TRUE)) +
  facet_wrap(vars(party_id), labeller = label_both)

Explore the fit

The code for that plot

References

Adolph, Christopher. 2025. “Problem Set 4: Modeling Presidential Approval with Ordered Probit.” POLS/CSSS 510: Maximum Likelihood Methods for the Social Sciences, University of Washington, Fall Quarter 2025. https://faculty.washington.edu/cadolph/mle/510hw4.pdf.
Miller, Warren E., Donald R. Kinder, Steven J. Rosenstone, and University of Michigan. Institute for Social Research. American National Election Studies. 2016. “ANES 1992 Time Series Study.” Inter-university Consortium for Political; Social Research [distributor]. https://doi.org/10.3886/ICPSR06067.v3.