A box plot summarizes a numerical distribution using its median, quartiles and whiskers. Place several boxes on the same scale to compare groups' centers and spread. Read the whisker rule before interpreting isolated points as potential outliers.
What does a box plot show?
Imagine sorting a set of test scores from lowest to highest. A box plot divides that ordered data into quarters. Its box spans the first quartile (Q1) to the third quartile (Q3), and the line inside marks the median.
| Part | What it tells you |
|---|---|
| Q1 | The value at the lower edge of the box |
| Median (Q2) | The line inside the box |
| Q3 | The value at the upper edge of the box |
| IQR | The box's width or height: Q3 − Q1 |
| Whiskers | The most extreme observed values inside the selected limits |
| Separate points | Observations outside those limits and worth investigating |
The distance from Q1 to Q3 is called the interquartile range, or IQR. You do not need to calculate it by hand to read the chart, but understanding it helps explain how the box and whiskers are constructed. A larger IQR means the central part of the distribution spans a wider range.
A common box plot convention calculates the interquartile range as IQR = Q3 − Q1. It then sets a lower fence at Q1 − 1.5 × IQR and an upper fence at Q3 + 1.5 × IQR. The fences are thresholds, not data points. Each whisker reaches the most extreme observed value inside its fence; observations beyond the fences are potential outliers and may be shown separately.
How to read a box plot, with examples
Read the box plot marks in this order:
- Locate the median. It shows the center of the ordered values, not their arithmetic mean.
- Read Q1 and Q3. Their distance is the IQR, a compact measure of spread.
- Check the whisker rule and endpoints. Do not assume that whiskers always equal the dataset's minimum and maximum.
- Investigate separate points. They are observations flagged by the chart's rule, not automatic errors.
- Compare groups on the same scale. Look at center and spread together, and consider the number of observations behind each box.
Box plot example: calculate quartiles and outliers
Imagine ten customers contacting a support team. These made-up values show how many minutes each customer waited for a first response, arranged from shortest to longest:
2, 4, 4, 4, 8, 10, 14, 14, 16, 30 minutes
For this example, we find Q1 and Q3 by taking the median of each sorted half. The lower half is 2, 4, 4, 4, 8, whose middle value is 4. The upper half is 10, 14, 14, 16, 30, whose middle value is 14. The median of all ten values is halfway between 8 and 10, giving 9 minutes.
Different software can calculate quartiles differently, so keep the method with the result.
| Calculation | Result |
|---|---|
| Q1 | 4 minutes |
| Median | (8 + 10) / 2 = 9 minutes |
| Q3 | 14 minutes |
| IQR | 14 − 4 = 10 minutes |
| Lower fence | 4 − 1.5 × 10 = −11 minutes |
| Upper fence | 14 + 1.5 × 10 = 29 minutes |
| Observed whiskers | 2 and 16 minutes |
| Potential outlier | 30 minutes |
The box spans Q1 = 4 to Q3 = 14 minutes, while one customer waited 30 minutes. That longer wait appears separately because it exceeds the calculated upper limit of 29 minutes. The upper whisker stops at 16 minutes: it must end at an actual observation within the limit, not at the limit itself. The negative lower limit is only a calculation; it does not mean anyone had a negative waiting time.
You can create a box-and-whisker plot from these values and compare the result with the calculation table.
Does an outlier mean the data are wrong?
No. An outlier is a value flagged by the plotting rule. The 30-minute wait could come from an unusually complicated question, a message sent outside working hours, or a recording problem. The box plot identifies an observation to investigate; the source data and context help explain it.
When to use a box plot
Use a box plot when you need to compare a numerical measure across groups, especially when many observations would make a dot-by-dot display crowded. Examples include response time by service, final grade by course, or compensation by role. Box plots work best when every group uses the same numerical scale.
A box plot is less useful when:
- every individual value matters in a small sample;
- the number of observations in each group must be immediately visible;
- the distribution's detailed shape, including multiple peaks or gaps, is central to the question.
For a small sample, a dot plot maker can keep every value visible. For a larger sample, you can inspect the distribution with a histogram. Pairing a box plot with either view can expose details hidden by the summary. A violin plot is another option when you want a compact view of distribution shape across groups.
Box plot vs. violin plot vs. histogram
| Chart | Best for | What it can hide |
|---|---|---|
| Box plot | Compact comparison of center and spread across groups | Sample size, clusters, gaps, and multiple peaks |
| Histogram | The shape of one numerical distribution | Exact observations; several groups can become cluttered |
| Violin plot | Comparing distribution shape across groups | Exact observations; unfamiliarity can slow interpretation |
Two datasets can share the same median and quartiles while having very different shapes. That is why a box plot is often a starting point rather than the entire analysis. Box plots become easier to judge when captions include sample sizes. Show the observations or pair the summary with a distribution chart when those details affect the conclusion.
Box plot examples with real data
The following examples use public datasets analyzed in Deepnote and report the results of completed analyses. They are separate from the illustrative response-time example above.
Student final grades
The UCI Student Performance dataset includes final grades for Mathematics and Portuguese, scored from 0 to 20. This example compares 395 Math records with 649 Portuguese records. Matching the identifying attributes documented with the dataset finds 382 students in both files, so many of the same students appear in both boxes.
Portuguese grades have a slightly higher median, and their IQR is narrower than the Math grades. The upper edge of both boxes is 14, while the Math box extends farther toward lower grades. That edge is the third quartile, not the highest grade in the dataset.
The comparison therefore shows more than a difference in typical performance: it also shows how widely the central grades vary within each course.
Video-delay A/B test
The public PlayDelay dataset comes from a Netflix A/B test in which devices were randomly assigned to a software version. It measures the delay between requesting playback and the video starting. The dataset records these time intervals to help researchers monitor streaming performance and detect regressions. Lower values mean shorter waits.
The delays are normalized, so they are not in seconds or milliseconds. The small decimal values show relative waiting times. Both groups use the same scale, so their positions and spreads can be compared. Each group contains about 343,000 observations. At that size, individual points beyond the whiskers would cover the chart, so this plot shows only the boxes and whiskers.
The tested version has a slightly longer median delay, and its IQR is slightly wider. On these two measures, it does not show the shorter, more consistent waits you would hope for. The box plot makes that comparison visible, but deciding whether the difference is meaningful requires the experiment’s statistical analysis.
Developer compensation
The 2025 Stack Overflow Developer Survey includes respondents’ roles and reported annual compensation, converted to US dollars. This example compares four roles using responses with a stated role and a compensation amount greater than zero. Respondents could select several roles, so one person’s compensation may appear in more than one group.
The groups range from 7,431 full-stack developer responses to 316 senior executive responses, so the senior executive box rests on far fewer answers. Like the A/B chart, this plot shows boxes and whiskers without individual points.
Engineering managers and senior executives have higher median compensation than the two developer groups, but compensation also varies within each role. The median gives a useful point of comparison; it is not an amount that everyone in that role earns. The boxes add context by showing the range from Q1 to Q3 for each role.
How to make a box plot
To create a box plot, start with one numerical observation per row. Add a category column if you want to compare groups. Keep the raw values rather than supplying only pre-calculated averages.
Before creating the box plot, decide and document:
- which field supplies the numerical values;
- which field, if any, defines the groups;
- how missing observations are handled;
- which quartile and whisker convention the software uses; and
- whether individual observations or only potential outliers should appear.
Make a box plot in Deepnote
Deepnote's box plot maker accepts uploaded or pasted data and a natural-language instruction. For example:
Create a box-and-whisker plot of these support waiting times in minutes: 2, 4, 4, 4, 8, 10, 14, 14, 16, 30. Use the median of each sorted half for Q1 and Q3 and the 1.5 × IQR rule for whiskers. Show potential outliers, label the axis in minutes, and include the method in the caption.
When two tools disagree on quartiles or whisker endpoints, ask the Deepnote Agent to calculate them with each method using the same source values. The results can then show whether the difference comes from the calculation method or from the data being compared.
For reproducibility, keep the source values, cleaning steps, quartile convention, code, and caption in the notebook. Deepnote’s run snapshots preserve the executed blocks, outputs, and execution metadata, while version history and review tools make changes to the method visible.
This reflects the position in our notebook manifesto: the notebook should hold both the analytical work and the record of how its result was produced.
Teams working locally can use deepnote sync to mirror a workspace into one .deepnote file per notebook and push local edits back. If the cloud copy has changed since you pulled it, the command reports a conflict instead of overwriting it.
This local CLI workflow is separate from Deepnote’s Git integration. For repository-managed projects, Git Directory Sync keeps each notebook in its own .deepnote file in the repository. A change to the quartile calculation can then be reviewed alongside the analysis it affects.
If you have not chosen a format yet, the AI chart generator can help you explore an appropriate design.
Make a box-and-whisker plot in Excel
In Excel for Microsoft 365, select the source data. On Windows, choose Insert > Insert Statistic Chart > Box and Whisker. On Mac, open the Insert tab, select the statistical chart icon, and choose Box and Whisker. In Format Data Series, choose whether to show inner points, outlier points, and mean markers. Excel also offers inclusive- and exclusive-median quartile calculations. Record that setting when comparing the result with another tool.
Common mistakes
- Treating fences as whisker endpoints. In a box plot, fences are calculated thresholds. A whisker stops at the most extreme observed value still inside its fence.
- Assuming every implementation uses the same convention. Quartile and whisker rules vary. Name the software setting or method.
- Deleting every separate point. Investigate its source and context before deciding whether it is an error.
- Ignoring sample size. A narrow box from a small group and one from a large group do not carry the same amount of evidence.
- Overlooking hidden shape. Box plots with similar summaries can conceal different clusters or multiple peaks.
- Claiming significance or causation from medians. The chart describes distributions; inferential tests and study design answer different questions.