Anil Keshwani

Anil Keshwani

Senior Machine Learning Engineer
ai-coustics


The Probability Distributions

Table of contents
  1. The Poisson Family
  2. The Poisson is the Binomial as nnn tends to infinity
  3. The Geometric Distribution
  4. The Geometric Distribution describes the number of trials you wait for a success
  5. Derivation of the Variance of a Geometrically Distributed Random Variable
  6. The Exponential Distribution
  7. The Exponential Distribution from the Poisson
  8. Derivation given a Poisson process

This post introduces a few of the most commonly encountered families of probability distributions: the Poisson, which models rates, and the geometric distribution, which models waiting times for discrete stochastic processes and the exponential, which is a continuous analogue of the geometric relating closely to the Poisson family.

This is a prototype post on probability distributions that I hope to expand in future. It aims to cover the relationships between the distributions most commonly encountered in undergraduate statistics courses and that are very commonly employed for social and life science applications. For now, it contains a discussion of the Poisson, geometric and exponential families of distributions. The content and layout of this post may change as it is updated.

The Poisson Family

The following discusses the Poisson distribution, which models rates of independent variables in time and space.

The Poisson is the Binomial as nn tends to infinity

The Poisson distribution is the binomial when nn tends to infinity, n→∞n \rightarrow \infty, for npnp fixed, np=λnp = \lambda (i.e. also as p→0p \rightarrow 0).

Define λ=np\lambda = np

We have Y∼Binom(n,p)Y \sim Binom(n, p) when:

p(Y=y)=(nk)⋅pk⋅(1−p)n−kp(Y = y) = \binom{n}{k} \cdot p^k \cdot (1-p)^{n-k}

and Y∼Pois(n,p)Y \sim Pois(n, p) when:

p(Y=y)=e−λ⋅λyy!p(Y = y) = \frac{e^{-\lambda} \cdot \lambda^y}{y!}

This has the straightforward interpretation of being the case when the number of trials is large e.g. radioisotope decay can be modelled as Poisson distributed because even a small quantity of a metal contains a very large number of atoms.

From λ=np\lambda = np we have that p=λnp = \dfrac{\lambda}{n}. Substitute λn\dfrac{\lambda}{n} into the binomial density function and let n→∞n \rightarrow \infty:

lim⁡n→∞(nk)⋅(λn)k⋅(1−λn)n−k expand combinatorial term=lim⁡n→∞n!k!(n−k)!⋅(λn)k⋅(1−λn)n−k remove constant terms=1k!⋅λk lim⁡n→∞n!(n−k)!⋅1nk⋅(1−λn)n−k separate index of 1−λn=1k!⋅λk lim⁡n→∞n!(n−k)!⋅1nk⋅(1−λn)n⋅(1−λn)−k \begin{aligned} & \lim_{n \rightarrow \infty} \binom{n}{k} \cdot \left( \dfrac{\lambda}{n} \right) ^k \cdot \left(1-\dfrac{\lambda}{n} \right)^{n-k} && \ \text{expand combinatorial term} \\ \\ &= \lim_{n \rightarrow \infty} \dfrac{n!}{k!(n-k)!} \cdot \left( \dfrac{\lambda}{n} \right) ^k \cdot \left(1-\dfrac{\lambda}{n} \right)^{n-k} && \ \text{remove constant terms} \\ &= \dfrac{1}{k!} \cdot \lambda^k \ \lim_{n \rightarrow \infty} \dfrac{n!}{(n-k)!} \cdot \dfrac{1}{n^k} \cdot \left(1-\dfrac{\lambda}{n} \right)^{n-k} && \ \text{separate index of } 1 - \dfrac{\lambda}{n} \\ &= \dfrac{1}{k!} \cdot \lambda^k \ \lim_{n \rightarrow \infty} \dfrac{n!}{(n-k)!} \cdot \dfrac{1}{n^k} \cdot \left(1-\dfrac{\lambda}{n} \right)^{n} \cdot \left(1-\dfrac{\lambda}{n} \right)^{-k} && \ \text{} \\ \end{aligned}

We can consider this expression within the limit in three component parts:

1k!⋅λk lim⁡n→∞n!(n−k)!⋅1nk⏟1⋅(1−λn)n⏟2⋅(1−λn)−k⏟3\dfrac{1}{k!} \cdot \lambda^k \ \lim_{n \rightarrow \infty} \underbrace{\dfrac{n!}{(n-k)!} \cdot \dfrac{1}{n^k}}_{1} \cdot \underbrace{\left(1-\dfrac{\lambda}{n} \right)^{n}}_{2} \cdot \underbrace{\left(1-\dfrac{\lambda}{n} \right)^{-k}}_{3}

Looking at the first component, we can expand the factorials in the numerator and denominator whilst simultaneously keeping in mind that there are kk divisions by nn:

lim⁡n→∞n!(n−k)!⋅1nk=lim⁡n→∞n⋅(n−1)⋅⋯⋅2⋅1(n−k)⋅(n−k−1)⋅⋯⋅2⋅1⋅1nk k terms cancel=lim⁡n→∞n⋅(n−1)⋅⋯⋅(n−(k−2))⋅(n−(k−1))nk=lim⁡n→∞nn⋅(n−1)n⋅⋯⋅(n−(k−2))n⋅(n−(k−1))nn is in the top and bottom of each=1⋅1⋅⋯⋅1⋅1=1\begin{aligned} &\lim_{n \rightarrow \infty} \dfrac{n!}{(n-k)!} \cdot \dfrac{1}{n^k} \\ &= \lim_{n \rightarrow \infty} \dfrac{n \cdot (n-1) \cdot \dots \cdot 2 \cdot 1}{(n-k) \cdot (n-k-1) \cdot \dots \cdot 2 \cdot 1} \cdot \dfrac{1}{n^k} && \ k \text{ terms cancel} \\ &= \lim_{n \rightarrow \infty} \dfrac{n \cdot (n-1) \cdot \dots \cdot (n - (k-2)) \cdot (n - (k-1))}{n^k} && \\ &= \lim_{n \rightarrow \infty} \dfrac{n}{n} \cdot \dfrac{(n-1)}{n} \cdot \dots \cdot \dfrac{(n - (k-2))}{n} \cdot \dfrac{(n - (k-1))}{n} && \text{n is in the top and bottom of each} \\ &= 1 \cdot 1 \cdot \dots \cdot 1 \cdot 1 \\ &= 1 \end{aligned}

Now turning to the second term:

lim⁡n→∞(1−λn)n\lim_{n \rightarrow \infty} \left(1-\dfrac{\lambda}{n} \right)^{n}

Recall the definition of ee is:

e=lim⁡x→∞(1+1x)xe = \lim_{x \rightarrow \infty} \left(1 + \frac{1}{x}\right)^x

Set x=−nλx = -\dfrac{n}{\lambda}, which is also equal to −1p-\frac{1}{p}. Substitute this into the definition of ee:

lim⁡n→∞(1−λn)nreexpress −λn=lim⁡n→∞(1+1x)nreexpress index n=lim⁡n→∞(1+1x)−x⋅λ=lim⁡n→∞(1+1x)x⋅(−λ)e reemerges=e−λ\begin{aligned} &\lim_{n \rightarrow \infty} \left(1 - \frac{\lambda}{n}\right)^n && \text{reexpress } - \frac{\lambda}{n} \\ &= \lim_{n \rightarrow \infty} \left(1 + \frac{1}{x}\right)^n && \text{reexpress index}\ n \\ &= \lim_{n \rightarrow \infty} \left(1 + \frac{1}{x}\right)^{-x \cdot \lambda} = \lim_{n \rightarrow \infty} \left(1 + \frac{1}{x}\right)^{x \cdot (-\lambda)} && e \ \text{reemerges} \\ &= e^{-\lambda} \\ \end{aligned}

Note that this is legitimate as our construction, x=−nλx = -\frac{n}{\lambda}, means that xx approaches infinity as nn does.

Finally, the third term

lim⁡n→∞(1−λn)−k\lim_{n\rightarrow \infty} \left(1-\dfrac{\lambda}{n} \right)^{-k}

As nn becomes large, λn\dfrac{\lambda}{n} approaches 00 and the argument of the limit becomes 11:

lim⁡n→∞(1−λn)−k=1−k=11k=1\begin{aligned} &\lim_{n\rightarrow \infty} \left(1-\dfrac{\lambda}{n} \right)^{-k} \\ &= 1^{-k} \\ &= \dfrac{1}{1^k} \\ &= 1 \end{aligned}

So reconstituting the three original terms:

1k!⋅λk lim⁡n→∞n!(n−k)!⋅1nk⏟1⋅(1−λn)n⏟2⋅(1−λn)−k⏟3\dfrac{1}{k!} \cdot \lambda^k \ \lim_{n \rightarrow \infty} \underbrace{\dfrac{n!}{(n-k)!} \cdot \dfrac{1}{n^k}}_{1} \cdot \underbrace{\left(1-\dfrac{\lambda}{n} \right)^{n}}_{2} \cdot \underbrace{\left(1-\dfrac{\lambda}{n} \right)^{-k}}_{3}

Having taken the limits for large nn we now have:

1k!⋅λk (1)⏟1⋅e−λ⏟2⋅(1)⏟3=e−λ⋅λkk!\begin{aligned} &\dfrac{1}{k!} \cdot \lambda^k \ \underbrace{(1)}_{1} \cdot \underbrace{e^{-\lambda}}_{2} \cdot \underbrace{(1)}_{3} \\ \\ &= \dfrac{e^{-\lambda} \cdot \lambda^k}{k!} \end{aligned}

We have arrived at the familiar formula for the Poisson density from earlier.

The Geometric Distribution

The Geometric Distribution describes the number of trials you wait for a success

The geometric distribution (the number of Bernoulli trials you wait for a success) is a discrete precursor to the exponential. Its variance is derived using the derivative of the sum of a geometric series.

For binomially distributed random variables, the number of trials until the first success is distributed:

P(Y=y)=py−1⋅p for y=1,2,…P(Y=y) = p^{y-1} \cdot p \ \text{for } y=1,2, \dots

As such we have Y∼Geom(p)Y \sim Geom(p) with pp the per-trial probability of a success. We have E(Y)=1p\mathbb{E}(Y) = \dfrac{1}{p}, which is intuitive, and Var(Y)=1−pp2\mathbb{V}ar(Y) = \dfrac{1-p}{p^2}, which is derived via the first derivative of the sum of a Geometric Series (see next section).

The geometric distribution is memoryless, in the sense that

P(Y>n+k ∣ Y>k)=P(Y>n)P(Y > n + k \ | \ Y > k) = P(Y > n)

This can be interpreted as saying that irrespective of the past and past failures, you always have as long to wait in expectation as if you had just started. It is a simple consequence of the independence of the Bernoulli trials.

Derivation of the Variance of a Geometrically Distributed Random Variable

Prerequisites

As prerequisites, we have that the sum of a geometric series is:

g(r)=∑k=0∞ark=a+ar+ar2+ar3+⋯=a1−r=a(1−r)−1g(r)=\sum\limits_{k=0}^\infty ar^k=a+ar+ar^2+ar^3+\cdots=\dfrac{a}{1-r}=a(1-r)^{-1}

and consequently the first derivative with respect to rr is:

g′(r)=∑k=1∞akrk−1=0+a+2ar+3ar2+⋯=a(1−r)2=a(1−r)−2g'(r)=\sum\limits_{k=1}^\infty akr^{k-1}=0+a+2ar+3ar^2+\cdots=\dfrac{a}{(1-r)^2}=a(1-r)^{-2}

Derivation

From Var(Y)=E(Y2)−E(Y)2\mathbb{V}ar(Y) = \mathbb{E}(Y^2) - \mathbb{E}(Y)^2 we first add 0=E(Y)−E(Y)0 = \mathbb{E}(Y) - \mathbb{E}(Y)

Var(Y)=E(Y2)−E(Y)2add −E(Y)+E(Y)Var(Y)=E(Y2)−E(Y)+E(Y)−E(Y)2Expectation linear operatorVar(Y)=E(Y(Y−1))+E(Y)−E(Y)\begin{aligned} \mathbb{V}ar(Y) &= \mathbb{E}(Y^2) - \mathbb{E}(Y)^2 && \text{add } - \mathbb{E}(Y) + \mathbb{E}(Y) \\ \\ \mathbb{V}ar(Y) &= \mathbb{E}(Y^2) - \mathbb{E}(Y) + \mathbb{E}(Y) - \mathbb{E}(Y)^2 && \text{Expectation linear operator} \\ \\ \mathbb{V}ar(Y) &= \mathbb{E}(Y(Y-1)) + \mathbb{E}(Y) - \mathbb{E}(Y) \end{aligned}

Focusing just on the first term, E(Y(Y−1))\mathbb{E}(Y(Y-1)), we have

E(Y(Y−1))=∑y=1∞p⋅(1−p)y−1⋅y⋅(y−1)pull out (1−p)\begin{aligned} &\mathbb{E}(Y(Y-1)) \\ \\ &= \sum_{y=1}^{\infty} p \cdot (1-p)^{y-1} \cdot y \cdot (y-1) && \text{pull out } (1-p) \\ \\ \end{aligned}

This resembles the right hand side of the first derivative sum, s=2a(1−r)3s = \dfrac{2a}{(1-r)^3}, with a=p⋅(1−p)a=p\cdot(1-p) and r=1−pr=1-p, which gives as an alternative expression for this first term

2p(1−p)(1−(1−p))3=2p(1−p)p3=2(1−p)p2\dfrac{2p(1-p)}{(1-(1-p))^3} = \dfrac{2p(1-p)}{p^3} = \dfrac{2(1-p)}{p^2}

Putting this back as the first term into the expression for the variance gives

σ2=E(Y(Y−1))+E(Y)−E(Y)2=2(1−p)p2+1p−1p2=2(1−p)+p−1p2=1−pp2 ■\begin{aligned} \sigma^2 &= \mathbb{E}(Y(Y-1)) + \mathbb{E}(Y) - \mathbb{E}(Y)^2 \\ \\ &= \dfrac{2(1-p)}{p^2} + \dfrac{1}{p} - \dfrac{1}{p^2} \\ \\ &= \dfrac{2(1-p) + p - 1}{p^2} \\ \\ &= \dfrac{1-p}{p^2} \ \blacksquare \end{aligned}

This yields the expression for the variance of the geometric distribution, σ2=1−pp2\sigma^2 = \frac{1-p}{p^2}.

The Exponential Distribution

The following discusses the Exponential Distribution, which models the waiting times between Poisson-distributed events and is a continuous analogue of the geometric.

The Exponential Distribution from the Poisson

The exponential distribution describes the probability distribution of time, tt, waited for the first occurrence of a Poisson-distributed event, Y∼Pois(λ)Y \sim Pois(\lambda).

(The fact that tt models the time until the first event makes the exponential a special case of the Gamma distribution with k=1k=1. In other words, the exponential family of distributions is a proper subset of the gamma family of distributions.)

Derivation given a Poisson process

In the waiting period, no events occur making P(Y=0)P(Y = 0) a definition of the probability of waiting time being equal to 11 time unit and since the Poisson assumes independently distributed events, we can multiply this tt times to yield a probability distribution of waiting times over time.

Given Y∼Pois(λ)Y \sim Pois(\lambda)

P(Y=y)=e−λ⋅λyy!P(Y = y) = \dfrac{e^{-\lambda} \cdot \lambda^y}{y!}

For P(Y=0)P(Y = 0) we have

P(Y=0)=e−λ⋅λ00!=e−λP(Y = 0) = \dfrac{e^{-\lambda} \cdot \lambda^0}{0!} = e^{-\lambda}

The probability of tt of these waiting times occurring consecutively is tt such terms11 Though crucially, more than tt such waiting times could occur so this is a lower bound according to the underlying Poisson model.

(e−λ)t=e−λ⋅t(e^{-\lambda})^t = e^{-\lambda \cdot t}

We interpret this as the probability of the waiting time being at least duration tt

P(T≥t)=e−λ⋅tP(T \geq t) = e^{-\lambda \cdot t}

Accordingly, the exponential distribution function (CDF) is given by the complement to one

P(T≤t)=1−e−λ⋅tP(T \leq t) = 1 - e^{-\lambda \cdot t}

Probability distribution function

We are required to differentiate this to give the probability distribution function (PDF).

Taking 1−e−λ⋅t1 - e^{-\lambda \cdot t} and setting u=−λ⋅tu = -\lambda \cdot t

1−eu; u=−λ⋅t1 - e^u; \ u = -\lambda \cdot t ddu=−eu; dudt=−λ; ddt=λ⋅e−λ⋅t\frac{d}{du} = -e^u; \ \frac{du}{dt} = -\lambda; \ \frac{d}{dt} = \lambda \cdot e^{-\lambda \cdot t}

This yields the expression for the probability distribution of the exponential from the Poisson

P(T=t)=λ⋅e−λ⋅tP(T = t) = \lambda \cdot e^{-\lambda \cdot t}

Footnotes

  1. Though crucially, more than tt such waiting times could occur so this is a lower bound according to the underlying Poisson model. ↩