Generalized Linear Models (GLMs) Module

1) Generalized Linear Models: Overview

Generalized linear models (GLMs) extend ordinary linear models to response variables that do not follow a normal distribution. They are useful when the response is a count, a proportion, or a continuous measurement with a strongly non-normal distribution.

This module is a work in progress. It currently focuses on count data and introduces Poisson and negative binomial regression. Additional tutorials for other GLM types will be added later.

Common GLM response distributions

The response distribution should match the type of data being modeled. The table below gives common examples and the link function most often used with each model type.

Distribution Typical data Common link Example application
Gaussian (normal) Continuous measurements that are approximately symmetric Identity Body size or temperature measurements
Poisson Non-negative integer counts Log Number of individuals, species, or events
Negative binomial Non-negative integer counts with more variation than a Poisson model can accommodate Log Abundance counts that vary substantially among samples
Binomial Proportions, successes out of trials, or two-category outcomes Logit Survival, infection status, or germination
Gamma Positive continuous measurements, often right-skewed Log Biomass, growth rate, or time to an event when zero is not possible
Tweedie Non-negative continuous measurements, including many zeros Log Biomass or productivity measurements with many zero observations

The link function connects the predictors to the expected response. With the log link used for Poisson and negative binomial models, predictions remain positive and effects are multiplicative on the response scale. The tutorial introduces these ideas using plain-language interpretation rather than mathematical notation.

2) Count Data Tutorial: Poisson and Negative Binomial Regression

The first tutorial uses crab count data. The response is the number of male crabs associated with each female crab. Female width is a continuous predictor, and female color is a categorical predictor. The tutorial works through a complete applied workflow, from exploring the data to checking model fit, comparing Poisson and negative binomial models, and plotting predictions on the original response scale.

Before class

Watch the following 3 videos before working through the R tutorial.

  1. GLM Part 1: A New Perspective
    A conceptual introduction to why generalized linear models are useful.

  2. GLM Part 2: Count Regression
    An introduction to modelling count data.

  3. Poisson Regression in R
    An R tidyverse coding example working through a Poisson GLM analysis that also demonstrates useful data wrangling and plotting tasks.

The first 2 videos are part of a larger GLM video series. The full series includes additional material that may be useful as you continue learning about GLMs.

Also Here’s another (~20 min) video if want to hear another explanation of working with count data and GLMs: - Regression with Count Data: Poisson and Negative Binomial
A concise, non-R overview of Poisson, quasi-Poisson, and negative binomial regression, including overdispersion and underdispersion. It also briefly introduces more advanced count-data models for cases with too many zeros (zero-inflated models) or when zero counts are impossible (zero-truncated models). These extensions are beyond the scope of this tutorial but provide useful context for choosing among count-data models.

Tutorial files

Download the tutorial script into a folder before working through it:

What the tutorial covers

The tutorial asks you to apply the following workflow to your own data:

  1. Load and tidy the crab data, then explore the response and explanatory variables with summary statistics and ggplot2 graphics.
  2. Fit a Poisson regression model with a continuous predictor and interpret the model output.
  3. Use DHARMa simulations and diagnostic plots to assess whether the model describes the data adequately.
  4. Fit a negative binomial model when the count data are more variable than expected under a Poisson model.
  5. Compare the Poisson and negative binomial models using AICc, while also considering diagnostics and biological plausibility.
  6. Plot model predictions and 95% confidence intervals on the original response scale.
  7. Estimate and plot means for the light and dark color groups, including their contrast as a response ratio.
  8. Fit a full model with separate width relationships for the two colors and compare the alternative predictor structures.

Main analytical choices

Response distribution

The response is a count, so the first candidate model is Poisson regression. A Poisson model assumes that the mean and variance are closely related. Ecological count data often vary more than this assumption allows. When that happens, a negative binomial model can provide a better description of the extra variation.

Continuous and categorical predictors

Female width is a continuous predictor, so the tutorial first evaluates whether the expected number of male crabs changes with width. Female color is a categorical predictor, so the tutorial estimates and compares the expected counts for the two color groups.

The full model asks whether the relationship between female width and the number of males differs between light and dark females. In a plot, this produces a separate fitted curve for each color.

Model diagnostics

Watch these videos before or during the diagnostic sections:

The tutorial uses the DHARMa package to simulate residuals and evaluate model fit. The most important checks are whether the residuals show an overall departure from the expected distribution and whether the model has too much or too little variation relative to the data.

A significant DHARMa test is a signal to investigate, not an automatic reason to discard a model. Look at the diagnostic plots, consider which feature of the data is causing the signal, and decide whether the problem matters for the biological question. A model can be useful even when a test detects a small departure, particularly when the main conclusions are stable and no important pattern remains in the residual plots.

For count models, also consider whether the data contain more zeros than expected. If zero inflation is substantial, a zero-inflated or hurdle model may be appropriate, but those models are beyond the scope of this first tutorial.

And for more detail, read the DHARMa vignettes

Comparing models

AICc is used to compare models fitted to the same response variable and the same observations. Lower AICc indicates more support from the models being compared. AICc does not prove that one model is biologically correct, so use it together with diagnostics, effect estimates, plots, and knowledge of the sampling process.

Interpreting model predictions

The tutorial uses ggeffects to plot predictions for continuous predictors and their 95% confidence intervals on the response scale. It uses emmeans to estimate and compare means for the categorical predictor. These plots make it easier to describe expected counts rather than interpreting log-scale coefficients directly.

Because Poisson and negative binomial models use a log link, model coefficients and predictions are initially estimated on the log scale. The tutorial back-transforms them to the original count scale so they can be interpreted as expected numbers of males. Exponentiating a coefficient gives a multiplicative effect, or rate ratio, rather than an additive change in counts.

For the categorical comparison, the two group means are compared as a response ratio for the same reason. A response ratio of 1 indicates equal expected counts, whereas values above or below 1 indicate proportionally higher or lower counts in one group. The back-transformation also means that confidence intervals on the response scale are often asymmetric.

What to report

In a methods section, report the response variable and its sampling unit, describe the relevant distributions of the response variable (i.e., why you used a Poisson or negative binomial GLM), the predictors and interactions, and how model diagnostics and model comparisons were used. Include the relevant R packages and functions when this helps readers reproduce the analysis.

In a results section, emphasize estimates on the original response scale, 95% confidence intervals, and plots of the fitted relationships. Explain any important diagnostic signal and how it affected the analysis. Avoid reporting only a list of p-values or treating AICc as a substitute for biological interpretation.

3) Additional GLM tutorials planned

Future tutorials may introduce:

  • Binomial models for proportions and two-category responses.
  • Tweedie models for non-negative continuous responses such as biomass when many observations are zero.

Additional resources

Count-data and diagnostic resources

Online books and courses

Back to top