Charts help us compare things quickly, whether we are reading an AI benchmark, following employment figures, or reviewing a business report. Understanding the comparison takes a little more attention: what do the numbers measure, and does the presentation help us read them correctly?
Misleading graphs become easier to recognize once you know which choices to question. The examples below cover inconsistent bar lengths, mismatched scales, missing context, and conclusions that go beyond the data. They include published graphics from OpenAI and xAI, several historical cases, and two clearly marked demonstrations using public data.
Each example has a clearer view built in Deepnote, either correcting the display or supplying context the original comparison leaves out. With your own source values available, you can use the AI chart generator to explore a correction and check whether the original message still holds.
How to spot a misleading graph
Before accepting a chart’s headline, check three things:
- The measurement: What does each bar, point, or slice represent? Look for units, dates, the group being counted, and whether the values are totals, percentages, averages, or running totals.
- The display: Do the marks match their labels? Check the baseline, axis direction, date order, and whether comparable measurements use the same scale.
- The conclusion: Does the chart support what the headline says? Consider whether missing records, different test conditions, or a longer time period would change the comparison.
These checks help you explain what is wrong with bad graphs instead of judging them only by their appearance. Misleading charts can have accurate labels while leaving out information needed to interpret the comparison.
Misleading axes, scales, and date order
1. A truncated bar chart: Fox Business tax rates
A Fox Business graphic from July 2012 compared a top marginal income tax rate of 35% with a proposed rate of 39.6%. Its vertical axis began at 34%, leaving a short visible bar for the first rate and a much taller one for the second.
- What misleads: Readers compare bar lengths. Cutting off most of each bar makes the increase look much larger relative to the starting rate.
- What to change: Start the bars at zero and keep the percentage labels.
The difference remains 4.6 percentage points in both views. What changes is its apparent size. The zero-baseline chart preserves the increase without making the proposed rate look several times larger. These are the historical rates shown in the graphic, not a comparison of anyone’s total tax bill.
When you make a bar graph in Deepnote, check the baseline before adjusting colors and labels. Bar length is part of the calculation readers see.
2. Mismatched scales: Planned Parenthood services
A graphic shown at a 2015 congressional hearing used crossing arrows to compare Planned Parenthood’s abortion services with cancer screening and prevention services. The arrows suggested that the first category had overtaken the second, although the printed counts showed otherwise.
- What misleads: Both series count services, but their vertical positions do not follow a shared numerical scale.
- What to change: Plot the reported endpoints on the same count axis.
The corrected comparison still shows screening and prevention services falling and abortions increasing. It no longer shows a crossover. Those are different claims, and the distinction is visible as soon as the counts share a scale.
This redraw connects only the two reported years; it does not invent observations between them. The figures count services, which should not be read as numbers of individual patients.
For measurements with the same units, you can build a line graph with a shared scale. If the units differ, separate panels often make the comparison easier to follow.
Bad graphs with distorted shapes and proportions
3. Inconsistent bar lengths: OpenAI’s GPT-5 launch
During OpenAI’s August 2025 GPT-5 launch, a coding benchmark chart showed equal-height bars labeled 69.1 and 30.8. A segment labeled 52.8 was taller than both. The visual comparison contradicted the printed scores. OpenAI acknowledged errors in its launch graphics (naturally, we offered to help with the charting).
- What misleads: The bar heights suggest a different ranking and size of difference from the numbers.
- What to change: Draw each model and setting separately on one scale, starting at zero.
The redraw makes the two GPT-5 settings distinct. Its reported thinking score exceeds o3’s; its score without thinking does not. These are separate evaluation results, so adding their percentages would not produce a meaningful combined score.
4. Convenient histogram bins: Old Faithful waiting times
A histogram groups numerical observations into intervals called bins. The public Old Faithful dataset contains 272 observations, including waiting times until the next eruption. Here, it demonstrates how the same data can look different when the intervals change.
- What can mislead: Choosing one bin width because it supports a preferred description of the data.
- What to change: Compare several reasonable widths while keeping the observations and numerical range fixed.
Look for the concentrations of shorter and longer waits as you change the bins. Wider intervals give a coarser account of where observations fall. Their bars can also be taller simply because each interval covers more minutes, so compare the overall pattern rather than treating bar heights across settings as equivalent measurements.
You can explore bin widths with Deepnote’s histogram maker. This is a demonstration of a charting choice, not evidence that the dataset’s publishers produced misleading graphs.
Missing data and selective time frames
5. Unequal time periods: the White House jobs comparison
An October 2022 White House chart compared average monthly job gains across presidents. Biden’s bar covered less than two years, while the earlier presidents’ bars covered complete four- or eight-year terms. His period also began during the recovery from the sharp pandemic employment decline.
- What misleads: A single average makes these different periods look directly comparable and hides the fall and recovery behind them.
- What to change: Show monthly employment around the pandemic before interpreting the averages.
The redraw uses the Bureau of Labor Statistics’ total nonfarm payroll employment series, which counts payroll jobs and adjusts for recurring seasonal patterns.
The longer view shows that the recovery began before the change in administration and continued afterward. It helps explain why a recovery-period average can be unusually high. It does not measure how many jobs a president personally caused to exist.
This uses revised historical employment data, so it supplies context rather than reproducing the exact data available when the original chart was published.
6. Survivors-only data: a fund comparison
Survivorship bias means looking only at the funds, businesses, or products that remain at the end of a period. Records that disappeared can matter just as much to the question.
The year-end 2024 SPIVA U.S. Scorecard provides a useful real-data demonstration. About half of the domestic equity funds in its 15-year starting group survived to the end. SPIVA accounts for non-surviving funds; its report is not the misleading example.
- What would mislead: Treating only the surviving funds as representative of the original group.
- What to change: Keep track of the starting group and account for funds that closed or merged.
The missing portion is large enough that a survivors-only comparison would describe a substantially different group. This chart shows the extent of that omission. Calculating its effect on returns would require the funds’ performance records as well.
When a conclusion goes beyond the graph
7. Different test conditions: xAI’s Grok 3 benchmark
In xAI’s February 2025 release, the AIME 2025 mathematics chart shows Grok 3 Think reaching 93.3%. Its hover label also gives 77.3 for the main bar and identifies 93.3 as the result with cons@64, which the release describes as its highest level of computation at test time.
- What can mislead: Comparing only the longest bars hides the change in evaluation settings.
- What to change: Give the additional-computation result its own labeled row.
Against the displayed o3-mini high score of 86.5, the two Grok results fall on opposite sides. The setting matters to the ranking. This redraw exposes that distinction; it does not establish an equal-compute comparison or independently retest the models.
8. Cumulative sales read as accelerating growth: Apple’s iPhone figures
Apple’s September 2013 presentation included cumulative iPhone sales. A cumulative figure is a running total: each period’s sales are added to everything counted before it. That total rises whenever additional phones are sold, even if fewer are sold in the latest quarter.
- What can mislead: Reading a rising cumulative line as proof that sales are increasing every quarter.
- What to change: Show sales for each quarter alongside a clearly labeled running total.
Using Apple’s reported first three fiscal quarters of 2013, quarterly sales fall while the running total rises. Neither view is mathematically wrong. They answer different questions: how many phones sold during a quarter, and how many sold across the period so far.
There is another useful check here. Third-quarter sales were higher than in the same quarter a year earlier, despite being lower than in the preceding quarter. For a seasonal business, a quarter-to-quarter decline alone does not establish a lasting downward trend.
The running total in this redraw starts in fiscal Q1 2013. It explains the calculation using reported quarterly results; it is not a reconstruction of Apple’s lifetime-sales slide.
How to correct a misleading graph in Deepnote
A useful correction lets someone inspect the source values, see what changed, and understand why the new view answers the question more clearly. Deepnote keeps the data, calculations, chart, and explanation together, so those checks can remain part of the analysis.
Recreate the chart from its source values
Start with a data file, a published table, or verified labels from the original chart. Record the units, dates, and source. If an image does not contain enough information for an exact reconstruction, state what is missing instead of filling the gaps with estimates.
For the tax-rate example, paste this complete prompt into Deepnote Agent:
Create and run a Python chart using these values from a historical 2012 tax-rate comparison: existing top marginal income tax rate, 35%; proposed rate, 39.6%. Use two bars starting at zero. Label the axis “Top marginal income tax rate (%)” and label both bars with their values. Title it “Existing and proposed top marginal rates, 2012.” Calculate the difference in percentage points in a separate block. Keep the values and chart code visible, and identify this as a reconstruction of the reported rates, not a comparison of total tax bills.
Agent can write and run the code while keeping the calculation available for review. For prepared data and a quick first visualization, the AI chart generator offers another starting point.
Change one chart choice at a time
Once the values are checked, isolate the choice you want to test. For the tax-rate chart, change only the baseline. For a histogram, change only the bin width. Keeping the other settings fixed makes it easier to see what caused the difference in appearance.
You can ask Agent to revise a Python-generated chart through its code. In a native Chart block, use Chart AI to configure the available settings from a description, then inspect the chosen fields, aggregation, filters, and scales.
Correcting misleading data visualization sometimes requires a different chart altogether. A running total may need a companion quarterly view; overlapping survey answers may need separate bars instead of slices. Our guide to types of graphs and charts can help you choose a format that matches the question.
Share the correction and its sources
Add a short explanation of what changed and what the result still cannot establish. A corrected employment chart, for example, can restore the missing timeline without proving the effect of a president’s policies.
Use Deepnote’s block sharing and embedding to share the chart output where your workspace allows it. Readers can see the comparison, while your team can return to the data and calculations when a source changes or someone questions the interpretation.
Keep a copy of the source data used for publication. That makes it possible to reproduce the example even if a public dataset is revised later.