FR EN

PR-6: Normal Distribution

Module 3 · Probability

Section 1: Introduction

A student scores 78 on an exam with a class mean of 70 and a standard deviation of 8. What fraction of the class scored lower? You don’t need to survey every student. With one formula and a single table, the normal distribution gives you the answer in under a minute — and by the end of this lesson, so will you.

That formula works because of a pattern that 18th-century astronomers stumbled upon: every time they recorded many independent measurements of the same star and plotted the errors, the result was a perfect bell-shaped curve. Abraham de Moivre described this curve mathematically in 1733; Carl Friedrich Gauss later showed it was the inevitable shape of measurement error — which is why it is also called the Gaussian distribution. When many small, independent sources of variation combine, their total always tends toward this same bell shape.

Today, the normal distribution describes heights, exam scores, manufacturing tolerances, blood pressure readings, and the distribution of sample means (which you’ll see in the Inference module). Learning to work with the normal distribution is not optional in statistics — it is the core skill.

After this lesson, you will be able to:

By the end of this lesson, you will be able to:

  • Describe the properties of the normal distribution and apply the Empirical Rule (68–95–99.7).
  • Use the z-table to find , , and for the standard normal.
  • Standardize a normal variable using and find probabilities.
  • Solve inverse normal problems: find given a cumulative probability.
  • Recognize when the normal model is and is not appropriate.

Section 2: Prerequisites

What you need coming in — and why it matters today:

  • Random variable notation (PR-4): The notation reads “X is distributed as.” Today we write . You also need the CDF interpretation — as cumulative area — which is how the z-table is organized.
  • Mean and standard deviation (PR-4, DS-5): is the center; measures spread as average distance from the mean. In PR-5 you computed and for a binomial — the normal uses the same parameters, just with a continuous bell curve instead of discrete bars.
  • Empirical Rule (DS-5): About 68%, 95%, and 99.7% of values fall within 1, 2, and 3 standard deviations of the mean for approximately normal data. Today we derive this from the z-table directly.
  • Z-score (DS-5): tells you how many standard deviations lies above or below the mean. In DS-5 you used z-scores for comparison. Today they become the key to computing probabilities.

Quick check — can you recall these?

If and the distribution is roughly symmetric about , about 68% of values fall within which range of ?

Success Factor:

What changes in this lesson: In DS-5 you used z-scores to compare values across different distributions. Here the z-score becomes the passport to a single probability table — the z-table — that works for every normal distribution ever. One table, infinite distributions. The key is the standardization formula; everything else follows mechanically from it. The same tools can be used to approximate binomial probabilities with the normal distribution, and INF-1 will use the same z-table to reason about sample means.

Retrieval Warm-up — from earlier lessons

Heights of adult women in a city are approximately normally distributed with mean 163 cm and standard deviation 7 cm. A woman is 177 cm tall. Compute her z-score.

A binomial random variable has trials and success probability . What are the mean and standard deviation?

Section 3: Core Concepts

How this section is organized: Five concepts build in order. Read C1–C2 for the shape and table structure, then C3 for the conversion formula that links any normal distribution to the table. C4 reverses the process (finding values from probabilities), and C5 teaches when not to use the normal model.

  • C1: Shape and Empirical Rule — what the normal curve looks like and the 68–95–99.7 pattern
  • C2: The standard normal Z and the z-table — how to read three types of probabilities
  • C3: Standardizing any normal variable — the bridge formula
  • C4: Inverse normal — finding a value from a probability
  • C5: Recognizing when the normal model fits

C0 — Areas as Proportions of Populations

The normal distribution is often introduced as a mathematical curve. But it is more intuitive to think of it as a description of a population.

Suppose the heights of adult women in a large city follow a normal distribution with mean cm and standard deviation cm. Imagine measuring every one of the city’s 100,000 adult women.

RangeApprox. countProportion
Between 156 cm and 170 cm (within 1σ)68,270 women68.3%
Between 149 cm and 177 cm (within 2σ)95,450 women95.5%
Between 142 cm and 184 cm (within 3σ)99,730 women99.7%
Shorter than 149 cm (more than 2σ below)2,275 women2.3%
Taller than 184 cm (more than 3σ above)135 women0.13%

This is the key insight: probabilities are areas, and areas are proportions of populations. When we say , we mean that if we drew a random woman from this population, she has a 2.3% chance of being shorter than 149 cm — equivalently, 2.3% of the population is below 149 cm.

The z-table you will use in this lesson is simply a lookup table for these proportions. Every area under the standard normal curve corresponds to a proportion of a standard normal population — and, after standardization, to a proportion of any normal population.


C1 — Properties of the Normal Distribution

Connect to PR-4: a continuous random variable like height or test score cannot use a PMF — probabilities correspond to areas under a curve, not point masses. The normal distribution provides the specific curve.

Normal Distribution

A random variable follows a normal distribution, written , if its probability density function has a symmetric bell shape centered at with spread controlled by .

Key properties:

  • Symmetric about — the left and right halves are mirror images.
  • Mean = Median = Mode = (all three coincide at the center).
  • Tails are asymptotic — they approach but never touch the horizontal axis.
  • Total area under the curve = 1 (as required for any probability distribution).
  • Fully described by two parameters: (location) and (spread).

Empirical Rule (68–95–99.7): For any :

  • — about 68% of values within 1σ
  • — about 95% within 2σ
  • — about 99.7% within 3σ

Example: Heights of adult women follow approximately (mean 162 cm, σ = 6 cm). About 95% of women are between 150 cm and 174 cm (162 ± 12 = 162 ± 2×6).

The second parameter in is the variance , not the standard deviation . Writing when you mean (with ) is a common error. Always check: is the number after the comma or ? Convention throughout this course is .

μ = σ =
Figure 1: The Empirical Rule for a normal distribution. Set μ and σ, then pick a band: about 68% of values fall within ±1σ of the mean, 95% within ±2σ, and 99.7% within ±3σ — leaving only about 0.3% in the two tails beyond ±3σ. The endpoints μ ± kσ recompute live in data units.

C2 — Standard Normal Distribution and the Z-Table

Rather than computing a different area formula for every possible and , we use a single reference distribution.

We write to mean “Z follows the standard normal distribution” — the special case where and .

Standard Normal Distribution

has mean 0 and standard deviation 1.

The z-table gives the left-tail cumulative area: for any value .

From this single entry, we derive three probability types:

  1. Left-tail: — read directly from the table.
  2. Right-tail: — complement rule.
  3. Between two values: — subtract the smaller left-tail from the larger.

The z-table for this course is reproduced below. Rows give the ones and tenths digits of ; columns give the hundredths digit.

Table de la loi normale centrée réduite (Z)

Aire à gauche de z

$z$ 0.000.010.020.030.040.050.060.070.080.09
0.0 0.50000.50400.50800.51200.51600.51990.52390.52790.53190.5359
0.1 0.53980.54380.54780.55170.55570.55960.56360.56750.57140.5753
0.2 0.57930.58320.58710.59100.59480.59870.60260.60640.61030.6141
0.3 0.61790.62170.62550.62930.63310.63680.64060.64430.64800.6517
0.4 0.65540.65910.66280.66640.67000.67360.67720.68080.68440.6879
0.5 0.69150.69500.69850.70190.70540.70880.71230.71570.71900.7224
0.6 0.72570.72910.73240.73570.73890.74220.74540.74860.75170.7549
0.7 0.75800.76110.76420.76730.77040.77340.77640.77940.78230.7852
0.8 0.78810.79100.79390.79670.79950.80230.80510.80780.81060.8133
0.9 0.81590.81860.82120.82380.82640.82890.83150.83400.83650.8389
1.0 0.84130.84380.84610.84850.85080.85310.85540.85770.85990.8621
1.1 0.86430.86650.86860.87080.87290.87490.87700.87900.88100.8830
1.2 0.88490.88690.88880.89070.89250.89440.89620.89800.89970.9015
1.3 0.90320.90490.90660.90820.90990.91150.91310.91470.91620.9177
1.4 0.91920.92070.92220.92360.92510.92650.92790.92920.93060.9319
1.5 0.93320.93450.93570.93700.93820.93940.94060.94180.94290.9441
1.6 0.94520.94630.94740.94840.94950.95050.95150.95250.95350.9545
1.7 0.95540.95640.95730.95820.95910.95990.96080.96160.96250.9633
1.8 0.96410.96490.96560.96640.96710.96780.96860.96930.96990.9706
1.9 0.97130.97190.97260.97320.97380.97440.97500.97560.97610.9767
2.0 0.97720.97780.97830.97880.97930.97980.98030.98080.98120.9817
2.1 0.98210.98260.98300.98340.98380.98420.98460.98500.98540.9857
2.2 0.98610.98640.98680.98710.98750.98780.98810.98840.98870.9890
2.3 0.98930.98960.98980.99010.99040.99060.99090.99110.99130.9916
2.4 0.99180.99200.99220.99250.99270.99290.99310.99320.99340.9936
2.5 0.99380.99400.99410.99430.99450.99460.99480.99490.99510.9952
2.6 0.99530.99550.99560.99570.99590.99600.99610.99620.99630.9964
2.7 0.99650.99660.99670.99680.99690.99700.99710.99720.99730.9974
2.8 0.99740.99750.99760.99770.99770.99780.99790.99790.99800.9981
2.9 0.99810.99820.99820.99830.99840.99840.99850.99850.99860.9986
3.0 0.99870.99870.99870.99880.99880.99890.99890.99890.99900.9990

Three worked examples using the table above:

Type 1 — Left-tail: . Find row 1.2, column 0.03: table entry = 0.8907. Done.

Type 2 — Right-tail: .

Type 3 — Between: .

The z-table gives — a left-tail area. If you look up and get 0.9332, that means 93.32% of the distribution is below 1.50. The right-tail probability is . Using 0.9332 directly as a right-tail probability is the single most common z-table error.

When computing , you must subtract the two left-tail areas, not add them. . Adding them () would give a probability of 1 — clearly wrong. The intuition: you want the area in the middle strip, which is the big left-area minus the small left-area that you don’t want.

Try it interactively. Use the visualization below to shade normal curve areas. Enter any z-score to see the corresponding probability, or switch to a general normal distribution to practice standardization.

Figure 2: Interactive normal curve. Switch between Standard N(0,1) and General N(μ, σ²) using the tabs. In General mode, enter μ, σ, and x — the z-score and probability update live, showing how standardization works. Use the Shade buttons to explore left-tail, right-tail, and between-values probabilities. Right-tail and between modes show the calculation, not just the answer.

C3 — Standardizing a Normal Variable

Any can be converted to the standard normal scale using the standardization formula.

Standardization (Z-Score Transformation)

For , compute:

Then , and you can use the z-table exactly as in C2.

Interpretation: tells you how many standard deviations is above (positive) or below (negative) the mean.

Mini-example: IQ scores follow (mean 100, ). Find .

Standardize: .

Look up .

About 88.5% of people score below 118.

You must standardize before using the z-table. The z-table works only for . Looking up directly in a table designed for produces a nonsense answer. Every non-standard normal problem requires the conversion step first.

μ = σ = x =
Figure 3: The standardization transformation. Set μ and σ, then drag x. The orange marker on the X curve and the blue marker on the Z curve sit at the same horizontal position — the same fraction of the bell — confirming that z = (x − μ) / σ is an axis relabelling that leaves the shaded probability unchanged.

C4 — Finding Values Given a Probability (Inverse Normal)

Sometimes the question is reversed: given a probability, find the corresponding value. This is the percentile or quantile problem.

Inverse Normal Lookup

Given a cumulative probability , find such that :

  1. Look up in the body of the z-table to find (the z-score corresponding to the th quantile).
  2. Unstandardize:

We use (read “z-star”) to distinguish this critical value from an ordinary observed z-score.

Mini-example: Exam scores follow (). Find the 90th percentile.

Step 1: Find such that . Scanning the z-table body, the closest entry is 0.8997 at . So .

Step 2: Unstandardize: .

The 90th percentile score is approximately 80.24.

In the inverse problem, search the table body (the probability values) to find , then use the row/column labels to read . Students sometimes look up the probability in the margins (the z-values) instead of the body — this produces a random number unrelated to the problem. Remember: body → → unstandardize.

Figure 4: Inverse normal lookup. Set a target cumulative probability p (the percentile). The left tail shades to area p and the boundary slides to z*, the critical value whose left-tail area equals p — this is the body → z* reading the z-table requires. Then set μ and σ to watch the unstandardize step x = μ + z*·σ compute live. Choose p → find z* → unstandardize to x.

C5 — Recognizing Non-Normal Shapes

The normal model is powerful, but it is not universal. Applying it blindly to non-normal data produces incorrect probabilities.

When to Use (and Not Use) the Normal Model

Normal model is appropriate when:

  • Data are roughly symmetric and unimodal
  • Data are continuous (or nearly so)
  • A histogram shows a bell shape
  • The Empirical Rule holds approximately

Normal model is NOT appropriate when:

  • Data are strongly right-skewed or left-skewed (use skewed distributions)
  • Data are discrete counts (use binomial, Poisson, etc.)
  • Data have heavy tails or multiple modes
  • Data are bounded between 0 and 1 (proportions — use binomial or beta)
Figure: When is the normal model a good fit? A symmetric, unimodal bell (left) is well described by the normal distribution. A right-skewed shape (like incomes), a bimodal shape (two distinct groups), and a few separated discrete counts are not — the normal model would misdescribe their shape.

Looking ahead: the normal distribution can even approximate discrete binomial data when and . INF-1 covers how the Central Limit Theorem makes sample means approximately normal regardless of the original population’s shape — which is why the normal distribution appears throughout all of inferential statistics.

Income distributions are a classic example where the normal model fails badly: they are strongly right-skewed with a long upper tail. Using the normal model for income would dramatically underestimate the probability of very high incomes. Similarly, the number of car accidents per day is a count — binomial or Poisson, not normal.

Section 4: Worked Examples

Example 1 — Fully Worked: Three Probability Types for Z

Find (a) , (b) , and (c) .

(a) — left-tail, read directly.

I notice this is a left-tail query, so the z-table gives the answer immediately. Row 1.4, column 0.05: .

(b) — right-tail, use complement.

The table gives left-tail areas. I choose to use the complement: .

Look up : row , column 0.02 → .

.

Sanity check: Since , the right-tail area should be greater than 0.5. ✓ (0.7324 > 0.5)

(c) — between two values, subtract.

I need the area in the strip between and . The strategy is: big left-area minus small left-area.

(row 0.8, column 0.05)

(row , column 0.00)

.


Example 2 — Prediction Checkpoint: Standardize and Find Probability

IQ scores are modeled as (). Find .

Setup:

The question asks for a right-tail probability. I must standardize first: .

Since , the value 118 is above the mean. The right-tail probability must therefore be less than 0.50, so the complement of the table’s left-tail area will be needed.

Show Solution

Since , the value 118 is above the mean. So the right-tail probability is less than 0.50.

(from the z-table in Section 3).

.

Therefore — about 11.5% of people score above 118.

Sanity check: Since 118 is 1.2 standard deviations above the mean, the Empirical Rule says about 84% of scores are below . Being below 1.2σ from the right puts us between 84% and 97.5%, so a right-tail of about 11.5% is consistent. ✓


Example 3 — Minimally Scaffolded: Inverse Normal (90th Percentile)

Exam scores follow (). Find the 90th percentile.

Hint: For the inverse problem, search the z-table body for the target probability, read off , then unstandardize.

Show Solution

Step 1: The 90th percentile means .

Step 2: Search the z-table body for 0.9000 (or the closest entry). The entry 0.8997 appears at ; the entry 0.9015 appears at . The closer value is .

Step 3: Unstandardize: .

The 90th percentile is approximately 80.24. A score of about 80 separates the top 10% from the rest of the class.


Example 4 — Precision Comparison: Empirical Rule Application

Analysis to examine:

Heights of adult males follow approximately ( cm). A student is asked: “What fraction of men are between 158 cm and 182 cm?”

The student writes: “158 and 182 are each standard deviations from the mean, so by the Empirical Rule, 95% of men fall in this range.”

A second student says: “The interval is 158 to 182, which is , so 95.45% of men are in this range.”

Both students identified the correct interval: and , so the interval is indeed .

The first student says “95%” — this is the rounded Empirical Rule value and is acceptable for quick reasoning.

The second student says “95.45%” — this is the more precise value from the z-table: .

For exam purposes, either “95%” or “95.45%” (equivalently, ) will be accepted. When you have access to the z-table (as in all exercises in this lesson), use the precise value.

Lesson: The Empirical Rule gives quick estimates; the z-table gives exact probabilities. Know when to use each.

Section 5: Guided Practice

Note: The z-table in Section 3 is scrollable and reusable. Keep it open in a separate browser tab or scroll back to Section 3 as needed while working through these problems.

Problem 1 — Standard Normal Probabilities (C2)


Problem 2 — Standardize and Find Probability (C3)


Problem 3 — Inverse Normal (C4)


Problem 4 — Classify the Distribution (C5)

For each description below, select the best classification-and-justification statement.

The ages at which people get their first driver’s licence in Quebec, recorded for all new licence holders in a given year.

The blood pressure readings (in mmHg) of a large sample of healthy adults, selected randomly from a national health survey.

The number of customer complaints received per day at a call centre, recorded over 200 working days.

The weights of apples harvested from a single large orchard, measured in grams for a random sample of 500 apples.

The annual household incomes of all residents in a city.

Section 6: Independent Practice

Reminder: The z-table in Section 3 is scrollable and reusable. Open Section 3 in a separate tab or scroll back as needed.

Problem 1 — Standard Normal Probability (C2)


Problem 2 — Standardize and Find One-Tail Probability (C3)


Problem 3 — Between Two Values (C3)


Problem 4 — Inverse Normal (C4)


Problem 5 — Empirical Rule Application (C1)


Problem 6 — Find the Error


Problem 7 — Multi-Step Synthesis: Quality Control at a Bottling Plant


Problem 8 — Re-centering a Process


Problem 9 — Lower Percentile


Mixed Review — Retrieval from Earlier Lessons



Spaced Review — Empirical Rule


Normal-Model Check (C5)

Section 7: Mastery Check

Question 1 — Feynman Test

In your own words, explain why using only the symmetry of the normal distribution. Do not use the z-table or specific numbers — just the concept of symmetry. Aim for 200–500 characters.

0 / 500

Question 2 — Apply: Between Two Bounds


Question 3 — Apply: Percentile Value


Question 4 — Diagnose a z-Table Lookup


Self-Assessment

How confident do you feel about the normal distribution and z-table lookups?

Still confusedReady for the Boss Fight

Section 8: Boss Fight

Choose your path. Both require standardization, table lookup, and inverse normal reasoning.

⚙️ Path A: The Engineer

A manufacturing process produces bolts whose diameters are normally distributed. You must find the rejection rate and set a threshold that rejects exactly the most extreme bolts.

📊 Path B: The Analyst

A professor wants to set the grade cutoff for an A. You must find the minimum score, then determine what spread would be needed if the cutoff must be at least 80.

⚙️ Path A: The Engineer

A manufacturing process produces bolts whose diameters follow (mean 10.0 mm, mm). Bolts are rejected if their diameter falls outside the interval mm.

Task 1. Find the probability that a randomly selected bolt is accepted (diameter falls within ). Show all standardization steps.

Show Guidance for Task 1

Standardize both bounds:

About 86.64% of bolts are accepted.


Task 2. What is the rejection rate? How many bolts out of 5,000 produced per day would you expect to be rejected?

Show Guidance for Task 2

Rejection rate = , so about 13.36% of bolts are rejected.

Expected daily rejections: bolts.


Task 3. The engineering team wants to adjust the upper rejection threshold (keeping the lower threshold at 9.7 mm) so that only the top 5% of bolt diameters are rejected on the upper end. Find the new upper threshold diameter.

Show Guidance for Task 3

We want , i.e., .

Find such that . From the z-table: .

Unstandardize: mm.

The new upper threshold is approximately 10.33 mm.


Task 4. With the original bounds and , the process engineer proposes reducing to 0.1 mm (by upgrading equipment). How does this change the acceptance rate? Is the investment worthwhile if each 0.1 mm reduction in costs $50,000?

Show Guidance for Task 4

With :

Acceptance rate improves from 86.64% to 99.74%. Rejection rate drops from 13.36% to 0.26%.

Out of 5,000 bolts/day: rejections drop from 668 to only 13 bolts/day — a reduction of 655 bolts/day.

The cost-benefit analysis depends on the value of each rejected bolt and the production cost. If each rejected bolt costs more than 76 in waste, the investment pays off daily. For typical manufacturing, this precision upgrade is often worthwhile.

Reflection: Write two sentences explaining how reducing σ affects the rejection rate, and why the acceptance interval [9.7, 10.3] matters for this calculation.

0 / 500

📊 Path B: The Analyst

A professor wants the top 15% of students to earn an A. Class scores follow (mean 72, ).

Task 1. Find the minimum score required to earn an A. Show the full inverse normal procedure.

Show Guidance for Task 1

We want , i.e., .

Find such that . From the z-table: (entry 0.8508, closest to 0.85).

Unstandardize: .

The minimum A score is approximately 82.4 (or 83 rounded up to a whole number).


Task 2. A student scores 85. What percentile does this correspond to? Interpret your answer in plain language.

Show Guidance for Task 2

Standardize: .

.

The student is at approximately the 90th percentile — scoring higher than about 90.3% of the class.

Plain language: “A score of 85 puts this student in the top 10% of the class, comfortably within A territory.”


Task 3. The professor teaches a second section where the mean stays at 72 but the spread is unknown. She wants the minimum A score to be at least 80. What is the maximum value of that achieves this?

Show Guidance for Task 3

We need the 85th percentile to be at least 80:

The standard deviation must be at most approximately 7.69 for the A cutoff to be 80 or higher. If scores are more spread out (), the top 15% threshold rises above 80.


Task 4. The professor considers a different grading policy: award an A to any student more than 1 standard deviation above the mean (i.e., ). Under , what percentage of students earn an A under this policy? Is this more or less than the “top 15%” policy?

Show Guidance for Task 4

.

About 15.87% of students would earn an A — slightly more than the 15% policy.

This makes sense: the “top 15%” policy requires , which is slightly above 1.00. Setting the cutoff at exactly corresponds to the 84.13th percentile, giving 15.87% of students an A versus exactly 15% under the stricter policy.

Reflection: Explain in two sentences why reducing σ (making the class more homogeneous) would raise the minimum A score when the top 15% rule is used.

0 / 500

Section 9: Challenge Problems

Ready for more? These go beyond the lesson objectives.

Challenge 1 — Mirror Tails

Starting from the normal distribution’s symmetry, determine how the left-tail probability below relates to the left-tail probability below , for any .

Then verify numerically for : look up both and in the z-table and confirm the relationship holds.

Show Solution

Proof:

The standard normal is symmetric about 0: for all .

By this symmetry, the area to the left of equals the area to the right of :

Now apply the complement rule — the total area is 1:

(For a continuous distribution, , so .)

Combining: .

Numerical verification for :

From the z-table: .

.

Also from the z-table: . ✓

The values match. This also confirms the familiar fact that the critical values each cut off 2.5% in their respective tails — the basis for the 95% confidence interval (covered in INF-2).


Challenge 2 — Recovering a Distribution

A symmetric distribution has and . Assuming , find and .

Show Solution

Setting up the equations:

From : This means .

From the z-table, , so .

The standardization gives: Equation 1:

From : From the z-table, , so .

Equation 2:

Solving the system:

Add Equation 1 and Equation 2:

Substitute into Equation 1:

Answer: , . The distribution is .

Check: By symmetry, both tail probabilities should be equal (both 10%) when the distribution is symmetric about , and indeed . ✓

Section 10: Solutions Reference

View Full Solutions Page

The solutions page includes worked solutions for all Guided Practice problems, Independent Practice problems, Mastery Check questions, Boss Fight tasks, and Challenge Problems.


Quick-Reference Formula Card

FormulaPurposeWhen to use
Standardize to Before any z-table lookup for a non-standard normal
Right-tail from left-tailWhen the table gives left-tail but you need right-tail
Between two valuesSubtract smaller left-area from larger
Unstandardize (inverse normal)After finding from a given cumulative probability

Z-Table Reading Guide

  1. The z-table in Section 3 gives — always a left-tail area.
  2. Row = ones and tenths digit of (e.g., row 1.4 for ).
  3. Column = hundredths digit (e.g., column 0.05 for ).
  4. For inverse problems: search the table body for the target probability, then read the row and column labels to get .
  5. Symmetry shortcut: — you can always convert negative z-values using the positive half of the table.

Common Errors to Avoid