Skip to main content
Daily Math Minute

Unit 2: Exploring Two-Variable Data

Two-Way Tables

Summarizing and analyzing bivariate categorical data with two-way tables.

Advanced25 min lesson3 min readUpdated August 12, 2026Author not yet attributed

Reading Association from a Table of Counts

A survey of 200 students records both grade level (Junior/Senior) and whether they hold a part-time job. Before reading on: if grade level had nothing to do with holding a job, what would you expect the proportion of job-holders to look like among juniors compared to among seniors?

Job: YesJob: NoTotal
Junior3070100
Senior6040100
Total90110200
Two-way table: grade level by part-time job status (n = 200).

Definition — Marginal and Conditional Distributions

A marginal distribution summarizes one variable alone, using the table's row or column totals. A conditional distribution summarizes one variable within a single category of the other — e.g., the distribution of job status among juniors only. Comparing conditional distributions across categories reveals whether an association exists between the two variables.

Worked Example — Comparing Conditional Distributions

The marginal proportion with a job is 90/200 = 45%. But conditionally: among juniors, 30/100 = 30% have a job; among seniors, 60/100 = 60% do. These conditional proportions (30% vs. 60%) differ sharply from each other and from the overall 45% — strong evidence of an association between grade level and job status. If the two variables were unrelated, both conditional proportions would land close to the marginal 45%.

This comparison can also be framed through expected counts: if grade level and job status were truly independent, the Junior/Yes cell would be expected to hold (row total × column total)/grand total = (100×90)/200 = 45 students — the actual count, 30, is well below that. This same logic — comparing observed counts to what independence would predict — is exactly what a later unit formalizes into a significance test.

Tip

Always compare conditional distributions in the same direction (e.g., job status given grade level, not the reverse) across every category — comparing mismatched directions can make a real association look weaker or stronger than it is.

Common Mistakes

  • Comparing raw counts (30 vs. 60) directly instead of proportions within each group.

    Raw counts don't account for group size — comparing 30/100 to 60/100 (both out of equal-sized groups here) is what actually reveals the association; with unequal group sizes, comparing raw counts would be misleading regardless.

Key Takeaways

  • A two-way table organizes counts of two categorical variables observed on the same individuals.
  • Comparing conditional distributions across categories — not raw counts — reveals whether an association exists.
  • Expected counts under independence give a numerical benchmark: observed counts far from expected suggest association.

Summary

Two-way tables describe relationships between categorical variables. The next lesson turns to two quantitative variables, and the line that best summarizes their relationship.

Sign in to track your progress and mark this lesson complete.

Track your progress