Joint probability distribution

Last updated December 14, 2024

X

Y

p(X)

p(Y)

Many sample observations (black) are shown from a joint probability distribution. The marginal densities are shown as well (in blue and in red).

In probability theory, the joint probability distribution is the probability distribution of all possible pairs of outputs of two random variables that are defined on the same probability space. The joint distribution can just as well be considered for any given number of random variables. The joint distribution encodes the marginal distributions, i.e. the distributions of each of the individual random variables and the conditional probability distributions, which deal with how the outputs of one random variable are distributed when given information on the outputs of the other random variable(s).

Examples
Draws from an urn
Coin flips
Rolling a die
Marginal probability distribution
Joint cumulative distribution function
Joint density function or mass function
Discrete case
Continuous case
Mixed case
Additional properties
Joint distribution for independent variables
Joint distribution for conditionally dependent variables
Covariance
Correlation
Important named distributions
See also
References
External links

In the formal mathematical setup of measure theory, the joint distribution is given by the pushforward measure, by the map obtained by pairing together the given random variables of the sample space's probability measure.

In the case of real-valued random variables, the joint distribution, as a particular multivariate distribution, may be expressed by a multivariate cumulative distribution function, or by a multivariate probability density function together with a multivariate probability mass function. In the special case of continuous random variables, it is sufficient to consider probability density functions, and in the case of discrete random variables, it is sufficient to consider probability mass functions.

Examples

Draws from an urn

Each of two urns contains twice as many red balls as blue balls, and no others, and one ball is randomly selected from each urn, with the two draws independent of each other. Let $A$ and $B$ be discrete random variables associated with the outcomes of the draw from the first urn and second urn respectively. The probability of drawing a red ball from either of the urns is 2/3, and the probability of drawing a blue ball is 1/3. The joint probability distribution is presented in the following table:

	A=Red	A=Blue	P(B)
B=Red	(2/3)(2/3)=4/9	(1/3)(2/3)=2/9	4/9+2/9=2/3
B=Blue	(2/3)(1/3)=2/9	(1/3)(1/3)=1/9	2/9+1/9=1/3
P(A)	4/9+2/9=2/3	2/9+1/9=1/3

Each of the four inner cells shows the probability of a particular combination of results from the two draws; these probabilities are the joint distribution. In any one cell the probability of a particular combination occurring is (since the draws are independent) the product of the probability of the specified result for A and the probability of the specified result for B. The probabilities in these four cells sum to 1, as with all probability distributions.

Moreover, the final row and the final column give the marginal probability distribution for A and the marginal probability distribution for B respectively. For example, for A the first of these cells gives the sum of the probabilities for A being red, regardless of which possibility for B in the column above the cell occurs, as 2/3. Thus the marginal probability distribution for $A$ gives $A$ 's probabilities unconditional on $B$ , in a margin of the table.

Coin flips

Consider the flip of two fair coins; let $A$ and $B$ be discrete random variables associated with the outcomes of the first and second coin flips respectively. Each coin flip is a Bernoulli trial and has a Bernoulli distribution. If a coin displays "heads" then the associated random variable takes the value 1, and it takes the value 0 otherwise. The probability of each of these outcomes is 1/2, so the marginal (unconditional) density functions are

P(A)=1/2\quad {\text{for}}\quad A\in \{0,1\};

P(B)=1/2\quad {\text{for}}\quad B\in \{0,1\}.

The joint probability mass function of $A$ and $B$ defines probabilities for each pair of outcomes. All possible outcomes are

(A=0,B=0),(A=0,B=1),(A=1,B=0),(A=1,B=1).

Since each outcome is equally likely the joint probability mass function becomes

P(A,B)=1/4\quad {\text{for}}\quad A,B\in \{0,1\}.

Since the coin flips are independent, the joint probability mass function is the product of the marginals:

P(A,B)=P(A)P(B)\quad {\text{for}}\quad A,B\in \{0,1\}.

Rolling a die

Consider the roll of a fair die and let $A=1$ if the number is even (i.e. 2, 4, or 6) and $A=0$ otherwise. Furthermore, let $B=1$ if the number is prime (i.e. 2, 3, or 5) and $B=0$ otherwise.

	1	2	3	4	5	6
A	0	1	0	1	0	1
B	0	1	1	0	1	0

Then, the joint distribution of $A$ and $B$ , expressed as a probability mass function, is

\mathrm {P} (A=0,B=0)=P\{1\}={\frac {1}{6}},\quad \quad \mathrm {P} (A=1,B=0)=P\{4,6\}={\frac {2}{6}},

\mathrm {P} (A=0,B=1)=P\{3,5\}={\frac {2}{6}},\quad \quad \mathrm {P} (A=1,B=1)=P\{2\}={\frac {1}{6}}.

These probabilities necessarily sum to 1, since the probability of some combination of $A$ and $B$ occurring is 1.

Marginal probability distribution

If more than one random variable is defined in a random experiment, it is important to distinguish between the joint probability distribution of X and Y and the probability distribution of each variable individually. The individual probability distribution of a random variable is referred to as its marginal probability distribution. In general, the marginal probability distribution of X can be determined from the joint probability distribution of X and other random variables.

If the joint probability density function of random variable X and Y is $f_{X,Y}(x,y)$ , the marginal probability density function of X and Y, which defines the marginal distribution, is given by:

$f_{X}(x)=\int f_{X,Y}(x,y)\;dy$
$f_{Y}(y)=\int f_{X,Y}(x,y)\;dx$

where the first integral is over all points in the range of (X,Y) for which X=x and the second integral is over all points in the range of (X,Y) for which Y=y.^[1]

Joint cumulative distribution function

For a pair of random variables $X,Y$ , the joint cumulative distribution function (CDF) $F_{X,Y}$ is given by^[2]^{: p. 89}

F_{X,Y}(x,y)=\operatorname {P} (X\leq x,Y\leq y)

(Eq.1)

where the right-hand side represents the probability that the random variable $X$ takes on a value less than or equal to $x$ and that $Y$ takes on a value less than or equal to $y$ .

For $N$ random variables $X_{1},\ldots ,X_{N}$ , the joint CDF $F_{X_{1},\ldots ,X_{N}}$ is given by

F_{X_{1},\ldots ,X_{N}}(x_{1},\ldots ,x_{N})=\operatorname {P} (X_{1}\leq x_{1},\ldots ,X_{N}\leq x_{N})

(Eq.2)

Interpreting the $N$ random variables as a random vector $\mathbf {X} =(X_{1},\ldots ,X_{N})^{T}$ yields a shorter notation:

F_{\mathbf {X} }(\mathbf {x} )=\operatorname {P} (X_{1}\leq x_{1},\ldots ,X_{N}\leq x_{N})

Joint density function or mass function

Discrete case

The joint probability mass function of two discrete random variables $X,Y$ is:

p_{X,Y}(x,y)=\mathrm {P} (X=x\ \mathrm {and} \ Y=y)

(Eq.3)

or written in terms of conditional distributions

p_{X,Y}(x,y)=\mathrm {P} (Y=y\mid X=x)\cdot \mathrm {P} (X=x)=\mathrm {P} (X=x\mid Y=y)\cdot \mathrm {P} (Y=y)

where $\mathrm {P} (Y=y\mid X=x)$ is the probability of $Y=y$ given that $X=x$ .

The generalization of the preceding two-variable case is the joint probability distribution of $n\,$ discrete random variables $X_{1},X_{2},\dots ,X_{n}$ which is:

p_{X_{1},\ldots ,X_{n}}(x_{1},\ldots ,x_{n})=\mathrm {P} (X_{1}=x_{1}{\text{ and }}\dots {\text{ and }}X_{n}=x_{n})

(Eq.4)

or equivalently

{\begin{aligned}p_{X_{1},\ldots ,X_{n}}(x_{1},\ldots ,x_{n})&=\mathrm {P} (X_{1}=x_{1})\cdot \mathrm {P} (X_{2}=x_{2}\mid X_{1}=x_{1})\\&\cdot \mathrm {P} (X_{3}=x_{3}\mid X_{1}=x_{1},X_{2}=x_{2})\\&\dots \\&\cdot P(X_{n}=x_{n}\mid X_{1}=x_{1},X_{2}=x_{2},\dots ,X_{n-1}=x_{n-1}).\end{aligned}}

.

This identity is known as the chain rule of probability.

Since these are probabilities, in the two-variable case

\sum _{i}\sum _{j}\mathrm {P} (X=x_{i}\ \mathrm {and} \ Y=y_{j})=1,\,

which generalizes for $n\,$ discrete random variables $X_{1},X_{2},\dots ,X_{n}$ to

\sum _{i}\sum _{j}\dots \sum _{k}\mathrm {P} (X_{1}=x_{1i},X_{2}=x_{2j},\dots ,X_{n}=x_{nk})=1.\;

Continuous case

The joint probability density function $f_{X,Y}(x,y)$ for two continuous random variables is defined as the derivative of the joint cumulative distribution function (see Eq.1 ):

f_{X,Y}(x,y)={\frac {\partial ^{2}F_{X,Y}(x,y)}{\partial x\partial y}}

(Eq.5)

This is equal to:

f_{X,Y}(x,y)=f_{Y\mid X}(y\mid x)f_{X}(x)=f_{X\mid Y}(x\mid y)f_{Y}(y)

where $f_{Y\mid X}(y\mid x)$ and $f_{X\mid Y}(x\mid y)$ are the conditional distributions of $Y$ given $X=x$ and of $X$ given $Y=y$ respectively, and $f_{X}(x)$ and $f_{Y}(y)$ are the marginal distributions for $X$ and $Y$ respectively.

The definition extends naturally to more than two random variables:

f_{X_{1},\ldots ,X_{n}}(x_{1},\ldots ,x_{n})={\frac {\partial ^{n}F_{X_{1},\ldots ,X_{n}}(x_{1},\ldots ,x_{n})}{\partial x_{1}\ldots \partial x_{n}}}

(Eq.6)

Again, since these are probability distributions, one has

\int _{x}\int _{y}f_{X,Y}(x,y)\;dy\;dx=1

respectively

\int _{x_{1}}\ldots \int _{x_{n}}f_{X_{1},\ldots ,X_{n}}(x_{1},\ldots ,x_{n})\;dx_{n}\ldots \;dx_{1}=1

Mixed case

The "mixed joint density" may be defined where one or more random variables are continuous and the other random variables are discrete. With one variable of each type

{\begin{aligned}f_{X,Y}(x,y)=f_{X\mid Y}(x\mid y)\mathrm {P} (Y=y)=\mathrm {P} (Y=y\mid X=x)f_{X}(x).\end{aligned}}

One example of a situation in which one may wish to find the cumulative distribution of one random variable which is continuous and another random variable which is discrete arises when one wishes to use a logistic regression in predicting the probability of a binary outcome Y conditional on the value of a continuously distributed outcome $X$ . One must use the "mixed" joint density when finding the cumulative distribution of this binary outcome because the input variables $(X,Y)$ were initially defined in such a way that one could not collectively assign it either a probability density function or a probability mass function. Formally, $f_{X,Y}(x,y)$ is the probability density function of $(X,Y)$ with respect to the product measure on the respective supports of $X$ and $Y$ . Either of these two decompositions can then be used to recover the joint cumulative distribution function:

{\begin{aligned}F_{X,Y}(x,y)&=\sum \limits _{t\leq y}\int _{s=-\infty }^{x}f_{X,Y}(s,t)\;ds.\end{aligned}}

The definition generalizes to a mixture of arbitrary numbers of discrete and continuous random variables.

Additional properties

Joint distribution for independent variables

In general two random variables $X$ and $Y$ are independent if and only if the joint cumulative distribution function satisfies

F_{X,Y}(x,y)=F_{X}(x)\cdot F_{Y}(y)

Two discrete random variables $X$ and $Y$ are independent if and only if the joint probability mass function satisfies

P(X=x\ {\mbox{and}}\ Y=y)=P(X=x)\cdot P(Y=y)

for all $x$ and $y$ .

While the number of independent random events grows, the related joint probability value decreases rapidly to zero, according to a negative exponential law.

Similarly, two absolutely continuous random variables are independent if and only if

f_{X,Y}(x,y)=f_{X}(x)\cdot f_{Y}(y)

for all $x$ and $y$ . This means that acquiring any information about the value of one or more of the random variables leads to a conditional distribution of any other variable that is identical to its unconditional (marginal) distribution; thus no variable provides any information about any other variable.

Joint distribution for conditionally dependent variables

If a subset $A$ of the variables $X_{1},\cdots ,X_{n}$ is conditionally dependent given another subset $B$ of these variables, then the probability mass function of the joint distribution is $\mathrm {P} (X_{1},\ldots ,X_{n})$ . $\mathrm {P} (X_{1},\ldots ,X_{n})$ is equal to $P(B)\cdot P(A\mid B)$ . Therefore, it can be efficiently represented by the lower-dimensional probability distributions $P(B)$ and $P(A\mid B)$ . Such conditional independence relations can be represented with a Bayesian network or copula functions.

Covariance

When two or more random variables are defined on a probability space, it is useful to describe how they vary together; that is, it is useful to measure the relationship between the variables. A common measure of the relationship between two random variables is the covariance. Covariance is a measure of linear relationship between the random variables. If the relationship between the random variables is nonlinear, the covariance might not be sensitive to the relationship, which means, it does not relate the correlation between two variables.

The covariance between the random variables $X$ and $Y$ is^[3]

\operatorname {cov} (X,Y)=\sigma _{XY}=E[(X-\mu _{x})(Y-\mu _{y})]=E(XY)-\mu _{x}\mu _{y}.

Correlation

There is another measure of the relationship between two random variables that is often easier to interpret than the covariance.

The correlation just scales the covariance by the product of the standard deviation of each variable. Consequently, the correlation is a dimensionless quantity that can be used to compare the linear relationships between pairs of variables in different units. If the points in the joint probability distribution of X and Y that receive positive probability tend to fall along a line of positive (or negative) slope, ρ_XY is near +1 (or −1). If ρ_XY equals +1 or −1, it can be shown that the points in the joint probability distribution that receive positive probability fall exactly along a straight line. Two random variables with nonzero correlation are said to be correlated. Similar to covariance, the correlation is a measure of the linear relationship between random variables.

The correlation coefficient between the random variables $X$ and $Y$ is

\rho _{XY}={\frac {\operatorname {cov} (X,Y)}{\sqrt {V(X)V(Y)}}}={\frac {\sigma _{XY}}{\sigma _{X}\sigma _{Y}}}.

Important named distributions

Named joint distributions that arise frequently in statistics include the multivariate normal distribution, the multivariate stable distribution, the multinomial distribution, the negative multinomial distribution, the multivariate hypergeometric distribution, and the elliptical distribution.

Related Research Articles

In probability theory and statistics, the cumulative distribution function (CDF) of a real-valued random variable $, or just distribution function of, evaluated at, is the probability that will take a value less than or equal to .$

Independence is a fundamental notion in probability theory, as in statistics and the theory of stochastic processes. Two events are independent, statistically independent, or stochastically independent if, informally speaking, the occurrence of one does not affect the probability of occurrence of the other or, equivalently, does not affect the odds. Similarly, two random variables are independent if the realization of one does not affect the probability distribution of the other.

In probability theory, a probability density function (PDF), density function, or density of an absolutely continuous random variable, is a function whose value at any given sample in the sample space can be interpreted as providing a relative likelihood that the value of the random variable would be equal to that sample. Probability density is the probability per unit length, in other words, while the absolute likelihood for a continuous random variable to take on any particular value is 0, the value of the PDF at two different samples can be used to infer, in any particular draw of the random variable, how much more likely it is that the random variable would be close to one sample compared to the other sample.

In probability, and statistics, a multivariate random variable or random vector is a list or vector of mathematical variables each of whose value is unknown, either because the value has not yet occurred or because there is imperfect knowledge of its value. The individual variables in a random vector are grouped together because they are all part of a single mathematical system — often they represent different properties of an individual statistical unit. For example, while a given person has a specific age, height and weight, the representation of these features of an unspecified person from within a group would be a random vector. Normally each element of a random vector is a real number.

<span class="mw-page-title-main">Multivariate normal distribution</span> Generalization of the one-dimensional normal distribution to higher dimensions

In probability theory and statistics, the multivariate normal distribution, multivariate Gaussian distribution, or joint normal distribution is a generalization of the one-dimensional (univariate) normal distribution to higher dimensions. One definition is that a random vector is said to be k-variate normally distributed if every linear combination of its k components has a univariate normal distribution. Its importance derives mainly from the multivariate central limit theorem. The multivariate normal distribution is often used to describe, at least approximately, any set of (possibly) correlated real-valued random variables, each of which clusters around a mean value.

In statistics, correlation or dependence is any statistical relationship, whether causal or not, between two random variables or bivariate data. Although in the broadest sense, "correlation" may indicate any type of association, in statistics it usually refers to the degree to which a pair of variables are linearly related. Familiar examples of dependent phenomena include the correlation between the height of parents and their offspring, and the correlation between the price of a good and the quantity the consumers are willing to purchase, as it is depicted in the so-called demand curve.

In probability theory and statistics, covariance is a measure of the joint variability of two random variables.

In probability theory and statistics, two real-valued random variables, $,, are said to be uncorrelated if their covariance,, is zero. If two variables are uncorrelated, there is no linear relationship between them.$

<span class="mw-page-title-main">Covariance matrix</span> Measure of covariance of components of a random vector

In probability theory and statistics, a covariance matrix is a square matrix giving the covariance between each pair of elements of a given random vector.

In probability theory and statistics, a Gaussian process is a stochastic process, such that every finite collection of those random variables has a multivariate normal distribution. The distribution of a Gaussian process is the joint distribution of all those random variables, and as such, it is a distribution over functions with a continuous domain, e.g. time or space.

<span class="mw-page-title-main">Mutual information</span> Measure of dependence between two variables

In probability theory and information theory, the mutual information (MI) of two random variables is a measure of the mutual dependence between the two variables. More specifically, it quantifies the "amount of information" obtained about one random variable by observing the other random variable. The concept of mutual information is intimately linked to that of entropy of a random variable, a fundamental notion in information theory that quantifies the expected "amount of information" held in a random variable.

In probability theory and statistics, the marginal distribution of a subset of a collection of random variables is the probability distribution of the variables contained in the subset. It gives the probabilities of various values of the variables in the subset without reference to the values of the other variables. This contrasts with a conditional distribution, which gives the probabilities contingent upon the values of the other variables.

<span class="mw-page-title-main">Dirichlet distribution</span> Probability distribution

In probability and statistics, the Dirichlet distribution, often denoted $, is a family of continuous multivariate probability distributions parameterized by a vector of positive reals. It is a multivariate generalization of the beta distribution, hence its alternative name of multivariate beta distribution (MBD). Dirichlet distributions are commonly used as prior distributions in Bayesian statistics, and in fact, the Dirichlet distribution is the conjugate prior of the categorical distribution and multinomial distribution.$

In statistics and information theory, a maximum entropy probability distribution has entropy that is at least as great as that of all other members of a specified class of probability distributions. According to the principle of maximum entropy, if nothing is known about a distribution except that it belongs to a certain class, then the distribution with the largest entropy should be chosen as the least-informative default. The motivation is twofold: first, maximizing entropy minimizes the amount of prior information built into the distribution; second, many physical systems tend to move towards maximal entropy configurations over time.

This article discusses how information theory is related to measure theory.

In statistics, an exchangeable sequence of random variables is a sequence X₁, X₂, X₃, ... whose joint probability distribution does not change when the positions in the sequence in which finitely many of them appear are altered. In other words, the joint distribution is invariant to finite permutation. Thus, for example the sequences

In probability theory, the family of complex normal distributions, denoted $or, characterizes complex random variables whose real and imaginary parts are jointly normal. The complex normal family has three parameters: location parameter μ, covariance matrix, and the relation matrix . The standard complex normal is the univariate distribution with,, and .$

In statistics and in probability theory, distance correlation or distance covariance is a measure of dependence between two paired random vectors of arbitrary, not necessarily equal, dimension. The population distance correlation coefficient is zero if and only if the random vectors are independent. Thus, distance correlation measures both linear and nonlinear association between two random variables or random vectors. This is in contrast to Pearson's correlation, which can only detect linear association between two random variables.

In probability theory and statistics, a complex random vector is typically a tuple of complex-valued random variables, and generally is a random variable taking values in a vector space over the field of complex numbers. If $are complex-valued random variables, then the n -tuple is a complex random vector. Complex random variables can always be considered as pairs of real random vectors: their real and imaginary parts.$

Poisson-type random measures are a family of three random counting measures which are closed under restriction to a subspace, i.e. closed under thinning. They are the only distributions in the canonical non-negative power series family of distributions to possess this property and include the Poisson distribution, negative binomial distribution, and binomial distribution. The PT family of distributions is also known as the Katz family of distributions, the Panjer or (a,b,0) class of distributions and may be retrieved through the Conway–Maxwell–Poisson distribution.

References

↑ Montgomery, Douglas C. (19 November 2013). Applied statistics and probability for engineers. Runger, George C. (Sixth ed.). Hoboken, NJ. ISBN 978-1-118-53971-2. OCLC 861273897.{{cite book}}: CS1 maint: location missing publisher (link)
↑ Park,Kun Il (2018). Fundamentals of Probability and Stochastic Processes with Applications to Communications. Springer. ISBN 978-3-319-68074-3.
↑ Montgomery, Douglas C. (19 November 2013). Applied statistics and probability for engineers. Runger, George C. (Sixth ed.). Hoboken, NJ. ISBN 978-1-118-53971-2. OCLC 861273897.{{cite book}}: CS1 maint: location missing publisher (link)

External links

"Joint distribution", Encyclopedia of Mathematics , EMS Press, 2001 [1994]
"Multi-dimensional distribution", Encyclopedia of Mathematics , EMS Press, 2001 [1994]
A modern introduction to probability and statistics : understanding why and how. Dekking, Michel, 1946-. London: Springer. 2005. ISBN 978-1-85233-896-1. OCLC 262680588.
"Joint continuous density function". PlanetMath .
Mathworld: Joint Distribution Function

This page is based on this Wikipedia article
Text is available under the CC BY-SA 4.0 license; additional terms may apply.
Images, videos and audio are available under their respective licenses.

[1] Montgomery, Douglas C. (19 November 2013). Applied statistics and probability for engineers. Runger, George C. (Sixth ed.). Hoboken, NJ. ISBN 978-1-118-53971-2. OCLC 861273897.{{cite book}}: CS1 maint: location missing publisher (link)

[KunIlPark-2] Park,Kun Il (2018). Fundamentals of Probability and Stochastic Processes with Applications to Communications. Springer. ISBN 978-3-319-68074-3.

[3] Montgomery, Douglas C. (19 November 2013). Applied statistics and probability for engineers. Runger, George C. (Sixth ed.). Hoboken, NJ. ISBN 978-1-118-53971-2. OCLC 861273897.{{cite book}}: CS1 maint: location missing publisher (link)

[1]

[2]

[3]