Probability & Statistics

A visual introduction

‡ The University of Manchester

September 2026

About Me

Where to Find Me

How These Decks Work

Nine seminars, one idea each

The R runs here, in this page

No install · no account · nothing to download

The grey boxes below are real R. Press Run and the code executes inside your browser — there is no server, and nothing you type is sent anywhere.

Press Run a few times without changing anything. The mean comes out different every time.

Nine seminars One idea each, in the order we meet them. Nothing here needs the seminar before it to make sense.
Moving around Arrows or space step through. Esc shows every slide at once, M opens the contents.
Writing on it B opens a blackboard and C draws straight over a slide, so worked examples happen in front of you.

Describing Data I

Seminar 1 · Data visualisation

Which Picture for Which Data?

Tick the display that suits each set of data. The reasoning appears underneath.

Describing Data II

Seminar 2 · Descriptive statistics

Thirteen Salaries at a Small Firm

The Same Thirteen Salaries, in R

Seminar 2 · Descriptive statistics

Every summary you just watched is one line of R.

Change the last salary to 96 and run it again. The mode and the median count positions, so they do not move; the mean and the standard deviation measure distances, so they do.

A Bound that Holds for Any Shape

Seminar 2 · Chebyshev

Chebyshev promises at least \(1 - 1/k^2\) of any distribution lies within \(k\) standard deviations. No normality, no symmetry, no assumptions at all.

The bound holds, but loosely — that is the price of assuming nothing. Swap rexp(n, 1)^2 for rnorm(n) and the gap widens further: for a normal, \(k = 2\) captures 95%, not the 75% Chebyshev guarantees.

Basic Concepts

Seminar 3 · Conditional probability and Bayes

Probability of an Event

Count how many ways the event can happen, then how many things can happen at all. The probability is the ratio of the two.

\[P(\color{#93C54B}{A}) = \frac{\#\ \text{ways } \color{#93C54B}{A} \text{ can happen}}{\#\ \text{of things that can happen}} = \frac{|\color{#93C54B}{A}|}{|\Omega|}\]

Flip a coin twice

\(\Omega = \{hh,\ ht,\ th,\ tt\} \quad \color{#6E7681}{\text{4 in all}}\)

H H H T T H T T

At least one head

\(\color{#93C54B}{A} = \{hh,\ ht,\ th\} \qquad P(\color{#93C54B}{A}) = \tfrac{3}{4} = 75\%\)

No heads

\(\color{#325D88}{B} = \{tt\} \qquad P(\color{#325D88}{B}) = \tfrac{1}{4} = 25\%\)

Roll two dice

\(\Omega = \{11,\ 12,\ \ldots,\ 66\} \quad \color{#6E7681}{\text{36 in all}}\)

At least one five

\(\color{#93C54B}{A} = \{15, 25, 35, 45, 51, 52, 53, 54, 55, 56, 65\}\)

\(P(\color{#93C54B}{A}) = \tfrac{11}{36}\)

Or split on the first die

\(P(\color{#93C54B}{A}) = \tfrac{1}{6} + \tfrac{5}{6} \times \tfrac{1}{6} = \tfrac{6}{36} + \tfrac{5}{36} = \tfrac{11}{36}\)

A tempting answer is 6 + 6 = 12, six rolls with a five on the first die and six with a five on the second. Five and five is in both lists, so twelve counts it twice. The split on the right avoids that, because its second case needs the first die not to be a five.

Sample Spaces and Events

Roll two dice and every possible result is on the board. The sample space \(\Omega\) is that whole board, and an event is any part of it you can point at.

Unions, Intersections and Complements

Two events on the same board. \(\color{#93C54B}{A} \cup \color{#325D88}{B}\) is “either”, \(\color{#93C54B}{A} \cap \color{#325D88}{B}\) is “both”, and \(\color{#93C54B}{A}^{c}\) is everything left over.

The Same Picture, Drawn as Circles

A Venn diagram is the board with the grid rubbed out. The four counts are the ones you just read off it, and they still add to 36.

Set Notation for Probability

Counting gives \(P(\color{#93C54B}{A}) = |\color{#93C54B}{A}|/|\Omega|\) and \(P(\color{#93C54B}{A}^{c}) = 1 - P(\color{#93C54B}{A})\). Two rules do the rest.

Either · \(\color{#93C54B}{A} \cup \color{#325D88}{B}\)

\[\begin{array}{@{}l@{\;\;}c@{\;\;}l@{}} P(\color{#93C54B}{A} \cup \color{#325D88}{B}) & = & P(\color{#93C54B}{A}) + P(\color{#325D88}{B}) - P(\color{#93C54B}{A} \cap \color{#325D88}{B})\\[4pt] \phantom{P(\color{#93C54B}{A} \cap \color{#325D88}{B})} & = & P(\color{#93C54B}{A}) + P(\color{#325D88}{B}) \qquad \color{#6E7681}{\text{if mutually exclusive}} \end{array}\]

Both · \(\color{#93C54B}{A} \cap \color{#325D88}{B}\)

\[\begin{array}{@{}l@{\;\;}c@{\;\;}l@{}} P(\color{#93C54B}{A} \cap \color{#325D88}{B}) & = & P(\color{#93C54B}{A}) \times P(\color{#325D88}{B} \mid \color{#93C54B}{A})\\[4pt] \phantom{P(\color{#93C54B}{A} \cap \color{#325D88}{B})} & = & P(\color{#93C54B}{A}) \times P(\color{#325D88}{B}) \qquad \color{#6E7681}{\text{if independent}} \end{array}\]

Conditional Probabilities

Being told that \(\color{#325D88}{B}\) happened throws away every outcome outside it. What is left is a smaller sample space, and the question is how much of that is \(\color{#93C54B}{A}\).

Ω B A out of all of Ω B out of B alone

\[P(\color{#93C54B}{A} \mid \color{#325D88}{B}) = \frac{P(\color{#93C54B}{A} \cap \color{#325D88}{B})}{P(\color{#325D88}{B})}\]

Knowing it changes nothing \(\color{#325D88}{B}\) is “the second die is a five”

123456123456

Six rolls survive and one of them starts with a three, so \(P(\color{#93C54B}{A} \mid \color{#325D88}{B}) = \tfrac{1}{6}\), which is \(P(\color{#93C54B}{A})\) over again. The dice are independent.

Knowing it changes everything \(\color{#325D88}{C}\) is “the total is six”

123456123456

Now only five rolls survive and one of them starts with a three, so \(P(\color{#93C54B}{A} \mid \color{#325D88}{C}) = \tfrac{1}{5}\). The total tells you something about the first die.

Here \(\color{#93C54B}{A}\) is “the first die is a three”. Reversing the bar asks something else, and that is what Bayes is for.

Conditioning Narrows the Sample Space

Eight in ten office managers would buy the software (\(\color{#93C54B}{A}\)), and four in ten of those would also upgrade (\(\color{#325D88}{B}\)). Of the rest, one in ten would. Given \(\color{#325D88}{B}\), how likely is \(\color{#93C54B}{A}\)?

The Law of Total Probability

Cut \(\Omega\) into pieces that do not overlap and leave nothing out. Every outcome in \(\color{#93C54B}{A}\) then lies in exactly one piece, so the probability of \(\color{#93C54B}{A}\) is its slices added back up.

B₁ B₂ B₃ A Ω Three slices of A, one per piece

\[P(\color{#93C54B}{A}) = \sum_{i} P(\color{#93C54B}{A} \mid \color{#325D88}{B_i})\, P(\color{#325D88}{B_i})\]

Supplier 1 Half the parts, two per hundred faulty

\[0.50 \times 0.02 = 0.010\]

Supplier 2 Three parts in ten, four per hundred faulty

\[0.30 \times 0.04 = 0.012\]

Supplier 3 One part in five, seven per hundred faulty

\[0.20 \times 0.07 = 0.014\]

Adding the three routes gives \(P(\text{faulty}) = 0.036\). Ask instead which supplier sent the faulty part in your hand, and you are asking for Bayes.

Librarian or Farmer?

Steve is very shy and withdrawn, invariably helpful but with very little interest in people or in the world of reality. A meek and tidy soul, he has a need for order and structure, and a passion for detail.

Twenty Farmers for Every Librarian

The Formula is Just That Picture

Move the Prior, Watch the Answer Move

Seminar 3 · the base rate does the work

The evidence is identical every time — only the prior changes. 1/21 is the real ratio, one librarian for every twenty farmers; at 0.50 you get 0.80. The description never had an opinion about how common librarians are, so the base rate has to supply it.

Monty Hall Problem

A car sits behind one of three doors and a goat behind each of the others. You pick door 1. The host, who knows where the car is, opens door 3 on a goat and offers you the swap. Do you take it?

Combinatorics and the Factorial

There are \(n\) choices and you are taking \(r\) of them. Two questions settle which formula you need, whether you sample with replacement and whether the order matters.

\[n! = n \times (n-1) \times \cdots \times 2 \times 1 \qquad 0! = 1\]

With replacement Order matters

\[n^{\,r}\]

Ten coin flips in a row

\(2^{10} = 1{,}024\)

Without replacement Order matters Permutation · \(^{n}P_{r}\)

\[\dfrac{n!}{(n-r)!}\]

Five cards off the top of the deck

\(\begin{array}{@{}r@{\;}c@{\;}l@{}} \dfrac{52!}{47!} & = & 52 \times 51 \times 50 \times 49 \times 48 \\[3pt] & = & 311{,}875{,}200 \end{array}\)

Without replacement Order does not matter Combination · \(^{n}C_{r}\)

\[\dfrac{n!}{(n-r)!\,r!} = \dbinom{n}{r}\]

One hand, dealt either way

K Q 2 3 7K 2 Q 7 3

\(\dbinom{52}{5} = 2{,}598{,}960\)

The last two differ by \(r!\), the orders the \(r\) you picked could have come in. Those 311,875,200 runs collapse to 2,598,960 hands because one hand can be dealt \(5! = 120\) ways.

Probability Distributions

Seminar 4 · Binomial, Poisson, exponential, normal

Random Variables

A random variable is not really a variable. It is a rule that hands every outcome in the sample space a number, which is what lets us do arithmetic with chance.

\[X : \Omega \to \mathbb{R} \qquad\qquad P(X = x)\]

Discrete Values you can list

Roll two dice and let \(\color{#93C54B}{X}\) be the total. It is the number already written in every cell of the board.

3 + 42 + 56 + 1

Six of the thirty-six cells carry a 7, so

\(P(\color{#93C54B}{X} = 7) = \tfrac{6}{36}\)

Continuous Values filling a range

Wait at a bus stop and let \(\color{#325D88}{X}\) be the minutes until the next arrival. It can be 3, or 3.1, or 3.0004.

3.03.13.0004

No single value can carry any probability, so ask for a stretch instead

\(P(2 \le \color{#325D88}{X} \le 4)\)

Which kind you have decides the tool. A discrete variable gives every value its own probability and you add them up. A continuous one spreads probability along a line and you measure the area under a curve. The binomial and the Poisson ahead are discrete, the exponential and the normal are continuous.

A Distribution Is a Function

Hand \(P\) a value of \(X\) and it hands back the probability of that value. That function is the distribution, and writing it down once replaces all the counting.

\[\sum_{k} P(X = k) = 1 \qquad\qquad \int p(x)\,dx = 1\]

Flip a coin four times \(\Omega\) has sixteen outcomes, \(\color{#93C54B}{X}\) counts the heads

1/1604/1616/1624/1631/164number of heads

Why bother Three things the function buys you

One summary. Write \(P(\color{#93C54B}{X})\) once and every probability is already in it, with no sample space left to count.

A model. The distribution is a claim about the process, so data can test it and estimate the numbers inside it.

Arithmetic. Add the bars over a range, or measure the area under the curve, and you have \(P(a \le \color{#93C54B}{X} \le b)\).

The two examples ahead already have names. Heads in \(n\) flips is the binomial, and a height or a measurement error is the normal, written \(\color{#93C54B}{X} \sim \text{Binomial}\) and \(\color{#325D88}{X} \sim \text{Normal}\).

Which Seller Do You Trust?

Three sellers on the same marketplace, same item, same price. The only thing you have to go on is what previous buyers said about them. Which one do you buy from?

Laplace’s Rule of Succession

In 1774 Laplace said not to rate a seller \(k/n\). Credit one imaginary good review and one bad one, then rate them \((k+1)/(n+2)\).

Where the Binomial Comes From

Every review is an independent draw at one constant unknown rate \(S\). We cannot work back to it, only forwards. Suppose \(S\) and see what it predicts.

What If the Rate Were Different?

One guess at \(S\) gives one number. Turn the dial and the top plot reshapes, while the bottom traces that number against every \(S\) it could be.

The Binomial in Standard Notation

Everything so far, in the notation the worksheet uses. The rate we have been calling \(S\) is written \(\color{#F47C3C}{p}\); its complement is \(\color{#B94A48}{q} = 1 - \color{#F47C3C}{p}\).

\[P(X = \color{#93C54B}{k}) \;=\; \binom{n}{\color{#93C54B}{k}}\; \color{#F47C3C}{p}^{\,\color{#93C54B}{k}}\; \color{#B94A48}{q}^{\,\color{#B94A48}{n-k}} \qquad X \sim \operatorname{Binomial}(n, \color{#F47C3C}{p})\]

One review is a Bernoulli trial, a single attempt with two outcomes, scored \(1\) for good and \(0\) for bad.

\[P(B_i = 1) = \color{#F47C3C}{p}, \quad P(B_i = 0) = \color{#B94A48}{q}, \quad \operatorname{E}[B_i] = \color{#F47C3C}{p}, \quad \operatorname{Var}(B_i) = \color{#F47C3C}{p}\color{#B94A48}{q}\]

\[X \;=\; B_1 + B_2 + \cdots + B_n, \qquad B_i \sim \operatorname{Bernoulli}(\color{#F47C3C}{p})\]

The Rare Event Limit of a Binomial Distribution

Hold the average \(\color{#F47C3C}{\lambda} = n\,\color{#F47C3C}{p}\) still and push \(n\) up. Each trial gets rarer, but you expect the same number of events.

The Poisson in Standard Notation

Put \(\color{#F47C3C}{p} = \color{#F47C3C}{\lambda}/n\) into the binomial and let \(n\) run away.

\[P(X = \color{#93C54B}{k}) \;=\; \binom{n}{\color{#93C54B}{k}}\big(\tfrac{\color{#F47C3C}{\lambda}}{n}\big)^{\color{#93C54B}{k}}\big(1-\tfrac{\color{#F47C3C}{\lambda}}{n}\big)^{n-\color{#93C54B}{k}}\]

↓ Group the factors by whether they still mention \(n\)

\[=\;\frac{\color{#F47C3C}{\lambda}^{\,\color{#93C54B}{k}}}{\color{#93C54B}{k}!}\;\cdot\;\frac{n(n-1)\cdots(n-\color{#93C54B}{k}+1)}{n^{\color{#93C54B}{k}}}\;\cdot\;\big(1-\tfrac{\color{#F47C3C}{\lambda}}{n}\big)^{n}\;\cdot\;\big(1-\tfrac{\color{#F47C3C}{\lambda}}{n}\big)^{-\color{#93C54B}{k}}\]

↓ As \(n \to \infty\) those three go to \(1\), to \(e^{-\color{#F47C3C}{\lambda}}\), and to \(1\)

\[P(X = \color{#93C54B}{k}) \;=\; \frac{\color{#F47C3C}{\lambda}^{\,\color{#93C54B}{k}}\, e^{-\color{#F47C3C}{\lambda}}}{\color{#93C54B}{k}!} \qquad X \sim \operatorname{Poisson}(\color{#F47C3C}{\lambda})\]

Check the approximation

Seminar 4 · exact vs approximate

Three mistakes per 10,000 transactions; 10,000 audited. What is the chance of four or more? Compute it both ways.

Both read 0.3527681, so every digit R prints agrees and the approximation costs you nothing here. Now set n <- 20; p <- 0.15 (still \(np = 3\)) and watch it open up. That is what “large \(n\), small \(p\)” is protecting you from.

Time Between Poisson Events

Count arrivals, or time the wait. The bridge is the \(\color{#93C54B}{k = 0}\) bar.

The Exponential in Standard Notation

The wait beats \(t\) exactly when nothing arrives in \([0, t]\), so read the Poisson at \(\color{#93C54B}{k = 0}\).

\[P(T > t) \;=\; P\big(N(t) = \color{#93C54B}{0}\big), \qquad N(t) \sim \operatorname{Poisson}(\color{#F47C3C}{\lambda} t)\]

↓ Put \(\color{#93C54B}{k = 0}\) into the Poisson formula

\[=\;\frac{(\color{#F47C3C}{\lambda} t)^{\color{#93C54B}{0}}\, e^{-\color{#F47C3C}{\lambda} t}}{\color{#93C54B}{0}!} \;=\; e^{-\color{#F47C3C}{\lambda} t}\]

↓ That is the survival curve, so the density is the slope of \(1 - e^{-\color{#F47C3C}{\lambda} t}\)

\[f(t) \;=\; \color{#F47C3C}{\lambda}\, e^{-\color{#F47C3C}{\lambda} t} \qquad T \sim \operatorname{Exponential}(\color{#F47C3C}{\lambda})\]

The Limit of a Binomial Distribution for Large n

Hold \(\color{#F47C3C}{p}\) still and push \(n\) up, and the bars settle onto a bell.

The Galton Board

Galton’s 1889 machine: a coin flip at every peg, and a binomial built out of falling balls.

The Idea Behind the Central Limit Theorem

Let’s Try Some Simulations

Throw \(\color{#F47C3C}{n}\) of them and add them up, 2,500 times.

The Normal Distribution

That bell has a formula. \(\color{#B94A48}{\mu}\) slides it, \(\color{#F47C3C}{\sigma}\) stretches it, and the area under it is always 1.

\[f(x) \;=\; \frac{1}{\color{#F47C3C}{\sigma}\sqrt{2\pi}}\; e^{-\frac{1}{2}\left(\frac{x-\color{#B94A48}{\mu}}{\color{#F47C3C}{\sigma}}\right)^{2}}\]

Expected Value, Variance and Standard Deviation

The Law of Large Numbers, Intuitively

Unpacking the Normal Formula

Every piece of the formula has a job. The \(e\) and the square make the bell, \(\sigma\) sets its width, \(\mu\) moves it, and the constant in front keeps the area at 1.

Where the \(e\) in the Normal Comes From

Forget the binomial for a moment. Ask only that errors in two directions be independent and that no direction be special, and the bell has no choice about its shape.

Where the \(\pi\) in the Normal Comes From

Call \(C\) the area under the bell \(e^{-x^2}\). Build the round hill from two of them, then measure its volume in two different ways.

The Box Plot on a Normal Curve

A box plot is five numbers and no assumptions. Lay it over a normal curve and all five land in fixed places, which is where the \(1.5 \times \text{IQR}\) rule comes from.

−3σ −2σ −1σ μ +1σ +2σ +3σ Outliers Outliers IQR Q1 Q3 Q1 − 1.5 × IQR Q3 + 1.5 × IQR Median −2.698σ −0.6745σ 0.6745σ 2.698σ 24.65% 50% 24.65%

\(Q_1\) and \(Q_3\) sit at \(\mp 0.6745\sigma\), so the fences fall at \(\mp 2.698\sigma\) and only \(0.7\%\) of a normal lies beyond them. In a sample of a thousand, about seven points get flagged with nothing wrong.

Functions of a Random Variable

For \(\color{#93C54B}{Y} = g(X)\), work through the CDF and never the density. Ask which values of \(X\) make \(\color{#93C54B}{Y}\) small enough.

\[F_{\color{#93C54B}{Y}}(y) \;=\; P(a X + b \le y) \;=\; P\big(X \le \tfrac{y-b}{a}\big)\]

↓ The right-hand side is just \(X\)’s own CDF, read at a shifted point

\[F_{\color{#93C54B}{Y}}(y) \;=\; F_X\big(\tfrac{y-b}{a}\big)\]

↓ Differentiate in \(y\); the chain rule brings down the \(1/a\)

\[f_{\color{#93C54B}{Y}}(y) \;=\; \tfrac{1}{|a|}\, f_X\big(\tfrac{y-b}{a}\big) \qquad\Longrightarrow\qquad \color{#93C54B}{Y} \sim \operatorname{N}\big(a\mu + b,\; a^{2}\sigma^{2}\big)\]

The Chi-Square Distribution

Square \(k\) independent standard normals and add them up. The total is never negative, and it is the yardstick for sample variances, since \((n-1)s^2/\sigma^2 \sim \chi^2_{n-1}\).

The Student’s \(t\) Distribution

Swap the unknown \(\sigma\) for the sample’s \(s\) and the bell gets fatter tails. Formally \(T = Z/\sqrt{V/k}\) with \(V \sim \chi^2_k\), which is why the chi-square came first.

Relationships Among Probability Distributions

Each distribution is one step from another: a sum, a limit, a square or a ratio.

Time between events Z² Z/√(V/k)

Thank You

Some slides are based on the videos of 3Blue1Brown.

Slides created with Quarto and reveal.js.

R runs in the browser via webR and the quarto-webr extension; interactive plots use Observable JS.

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License. To view a copy of this license, visit:

https://creativecommons.org/licenses/by-sa/4.0/