Generalized Linear Models (GLMs) Module
1) Generalized Linear Models: Overview
Generalized linear models (GLMs) extend ordinary linear models to response variables that do not follow a normal distribution. They are useful when the response is a count, a proportion, or a continuous measurement with a strongly non-normal distribution.
This module is a work in progress. It currently focuses on count data and introduces Poisson and negative binomial regression. Additional tutorials for other GLM types will be added later.
Common GLM response distributions
The response distribution should match the type of data being modeled. The table below gives common examples and the link function most often used with each model type.
| Distribution | Typical data | Common link | Example application |
|---|---|---|---|
| Gaussian (normal) | Continuous measurements that are approximately symmetric | Identity | Body size or temperature measurements |
| Poisson | Non-negative integer counts | Log | Number of individuals, species, or events |
| Negative binomial | Non-negative integer counts with more variation than a Poisson model can accommodate | Log | Abundance counts that vary substantially among samples |
| Binomial | Proportions, successes out of trials, or two-category outcomes | Logit | Survival, infection status, or germination |
| Gamma | Positive continuous measurements, often right-skewed | Log | Biomass, growth rate, or time to an event when zero is not possible |
| Tweedie | Non-negative continuous measurements, including many zeros | Log | Biomass or productivity measurements with many zero observations |
The link function connects the predictors to the expected response. With the log link used for Poisson and negative binomial models, predictions remain positive and effects are multiplicative on the response scale. The tutorial introduces these ideas using plain-language interpretation rather than mathematical notation.
2) Count Data Tutorial: Poisson and Negative Binomial Regression
The first tutorial uses crab count data. The response is the number of male crabs associated with each female crab. Female width is a continuous predictor, and female color is a categorical predictor. The tutorial works through a complete applied workflow, from exploring the data to checking model fit, comparing Poisson and negative binomial models, and plotting predictions on the original response scale.
Before class
Watch the following 3 videos before working through the R tutorial.
GLM Part 1: A New Perspective
A conceptual introduction to why generalized linear models are useful.GLM Part 2: Count Regression
An introduction to modelling count data.Poisson Regression in R
An R tidyverse coding example working through a Poisson GLM analysis that also demonstrates useful data wrangling and plotting tasks.
The first 2 videos are part of a larger GLM video series. The full series includes additional material that may be useful as you continue learning about GLMs.
Also Here’s another (~20 min) video if want to hear another explanation of working with count data and GLMs: - Regression with Count Data: Poisson and Negative Binomial
A concise, non-R overview of Poisson, quasi-Poisson, and negative binomial regression, including overdispersion and underdispersion. It also briefly introduces more advanced count-data models for cases with too many zeros (zero-inflated models) or when zero counts are impossible (zero-truncated models). These extensions are beyond the scope of this tutorial but provide useful context for choosing among count-data models.
Tutorial files
Download the tutorial script into a folder before working through it:
What the tutorial covers
The tutorial asks you to apply the following workflow to your own data:
- Load and tidy the crab data, then explore the response and explanatory variables with summary statistics and
ggplot2graphics. - Fit a Poisson regression model with a continuous predictor and interpret the model output.
- Use DHARMa simulations and diagnostic plots to assess whether the model describes the data adequately.
- Fit a negative binomial model when the count data are more variable than expected under a Poisson model.
- Compare the Poisson and negative binomial models using AICc, while also considering diagnostics and biological plausibility.
- Plot model predictions and 95% confidence intervals on the original response scale.
- Estimate and plot means for the light and dark color groups, including their contrast as a response ratio.
- Fit a full model with separate width relationships for the two colors and compare the alternative predictor structures.
Main analytical choices
Response distribution
The response is a count, so the first candidate model is Poisson regression. A Poisson model assumes that the mean and variance are closely related. Ecological count data often vary more than this assumption allows. When that happens, a negative binomial model can provide a better description of the extra variation.
Link function
Poisson and negative binomial regression normally use a log link. This ensures that predicted counts are positive. Coefficients are estimated on the log scale, but predictions and effect estimates can be back-transformed to the original count scale for interpretation.
Because the back-transformation is nonlinear, confidence intervals on the response scale are often not symmetrical around the estimate. This is expected for count models and is one reason to show the confidence intervals in plots rather than relying only on coefficient tables.
Continuous and categorical predictors
Female width is a continuous predictor, so the tutorial first evaluates whether the expected number of male crabs changes with width. Female color is a categorical predictor, so the tutorial estimates and compares the expected counts for the two color groups.
The full model asks whether the relationship between female width and the number of males differs between light and dark females. In a plot, this produces a separate fitted curve for each color.
Model diagnostics
Watch these videos before or during the diagnostic sections:
The tutorial uses the DHARMa package to simulate residuals and evaluate model fit. The most important checks are whether the residuals show an overall departure from the expected distribution and whether the model has too much or too little variation relative to the data.
A significant DHARMa test is a signal to investigate, not an automatic reason to discard a model. Look at the diagnostic plots, consider which feature of the data is causing the signal, and decide whether the problem matters for the biological question. A model can be useful even when a test detects a small departure, particularly when the main conclusions are stable and no important pattern remains in the residual plots.
For count models, also consider whether the data contain more zeros than expected. If zero inflation is substantial, a zero-inflated or hurdle model may be appropriate, but those models are beyond the scope of this first tutorial.
And for more detail, read the DHARMa vignettes
Comparing models
AICc is used to compare models fitted to the same response variable and the same observations. Lower AICc indicates more support from the models being compared. AICc does not prove that one model is biologically correct, so use it together with diagnostics, effect estimates, plots, and knowledge of the sampling process.
Interpreting model predictions
The tutorial uses ggeffects to plot predictions for continuous predictors and their 95% confidence intervals on the response scale. It uses emmeans to estimate and compare means for the categorical predictor. These plots make it easier to describe expected counts rather than interpreting log-scale coefficients directly.
Because Poisson and negative binomial models use a log link, model coefficients and predictions are initially estimated on the log scale. The tutorial back-transforms them to the original count scale so they can be interpreted as expected numbers of males. Exponentiating a coefficient gives a multiplicative effect, or rate ratio, rather than an additive change in counts.
For the categorical comparison, the two group means are compared as a response ratio for the same reason. A response ratio of 1 indicates equal expected counts, whereas values above or below 1 indicate proportionally higher or lower counts in one group. The back-transformation also means that confidence intervals on the response scale are often asymmetric.
What to report
In a methods section, report the response variable and its sampling unit, describe the relevant distributions of the response variable (i.e., why you used a Poisson or negative binomial GLM), the predictors and interactions, and how model diagnostics and model comparisons were used. Include the relevant R packages and functions when this helps readers reproduce the analysis.
In a results section, emphasize estimates on the original response scale, 95% confidence intervals, and plots of the fitted relationships. Explain any important diagnostic signal and how it affected the analysis. Avoid reporting only a list of p-values or treating AICc as a substitute for biological interpretation.
3) Additional GLM tutorials planned
Future tutorials may introduce:
- Binomial models for proportions and two-category responses.
- Tweedie models for non-negative continuous responses such as biomass when many observations are zero.
Additional resources
Count-data and diagnostic resources
GLM Part 4: Overdispersion
Use the series link to locate the section on overdispersion.GLM Part 5: Diagnostics
Use the series link to locate the section on model diagnostics.(Video) Regression with Count Data: Poisson and Negative Binomial
A concise, non-R overview of Poisson, quasi-Poisson, and negative binomial regression, including overdispersion and underdispersion. It also briefly introduces more advanced count-data models for cases with too many zeros (zero-inflated models) or when zero counts are impossible (zero-truncated models). These extensions are beyond the scope of this tutorial but provide useful context for choosing among count-data models.
Online books and courses
- Frans Rodenburg, Applied Statistics: Generalized Linear Models
Background explanations and worked examples. Some examples use base R code. - QERM 514: Fitting count models
Examples of more advanced count-model options. - Zuur, Ieno, and Elphick, A Beginner’s Guide to GLM and GLMM with R
See this book for additional detail and extended case studies applying GLMs and GLMMs. Can also get it from the CPP Library