Ziggurat algorithm

Last updated August 29, 2023

The ziggurat algorithm is an algorithm for pseudo-random number sampling. Belonging to the class of rejection sampling algorithms, it relies on an underlying source of uniformly-distributed random numbers, typically from a pseudo-random number generator, as well as precomputed tables. The algorithm is used to generate values from a monotonically decreasing probability distribution. It can also be applied to symmetric unimodal distributions, such as the normal distribution, by choosing a value from one half of the distribution and then randomly choosing which half the value is considered to have been drawn from. It was developed by George Marsaglia and others in the 1960s.

A typical value produced by the algorithm only requires the generation of one random floating-point value and one random table index, followed by one table lookup, one multiply operation and one comparison. Sometimes (2.5% of the time, in the case of a normal or exponential distribution when using typical table sizes)^{[ citation needed ]} more computations are required. Nevertheless, the algorithm is computationally much faster^{[ citation needed ]} than the two most commonly used methods of generating normally distributed random numbers, the Marsaglia polar method and the Box–Muller transform, which require at least one logarithm and one square root calculation for each pair of generated values. However, since the ziggurat algorithm is more complex to implement it is best used when large quantities of random numbers are required.

The term ziggurat algorithm dates from Marsaglia's paper with Wai Wan Tsang in 2000; it is so named because it is conceptually based on covering the probability distribution with rectangular segments stacked in decreasing order of size, resulting in a figure that resembles a ziggurat.

Theory of operation

The ziggurat algorithm is a rejection sampling algorithm; it randomly generates a point in a distribution slightly larger than the desired distribution, then tests whether the generated point is inside the desired distribution. If not, it tries again. Given a random point underneath a probability density curve, its x coordinate is a random number with the desired distribution.

The distribution the ziggurat algorithm chooses from is made up of n equal-area regions; n − 1 rectangles that cover the bulk of the desired distribution, on top of a non-rectangular base that includes the tail of the distribution.

Given a monotone decreasing probability density function f(x), defined for all x ≥ 0, the base of the ziggurat is defined as all points inside the distribution and below y₁ = f(x₁). This consists of a rectangular region from (0, 0) to (x₁, y₁), and the (typically infinite) tail of the distribution, where x > x₁ (and y < y₁).

This layer (call it layer 0) has area A. On top of this, add a rectangular layer of width x₁ and height A/x₁, so it also has area A. The top of this layer is at height y₂ = y₁ + A/x₁, and intersects the density function at a point (x₂, y₂), where y₂ = f(x₂). This layer includes every point in the density function between y₁ and y₂, but (unlike the base layer) also includes points such as (x₁, y₂) which are not in the desired distribution.

Further layers are then stacked on top. To use a precomputed table of size n (n = 256 is typical), one chooses x₁ such that x_n = 0, meaning that the top box, layer n − 1, reaches the distribution's peak at (0, f(0)) exactly.

Layer i extends vertically from y_i to y_i+1, and can be divided into two regions horizontally: the (generally larger) portion from 0 to x_i+1 which is entirely contained within the desired distribution, and the (small) portion from x_i+1 to x_i, which is only partially contained.

Ignoring for a moment the problem of layer 0, and given uniform random variables U₀ and U₁ ∈ [0,1), the ziggurat algorithm can be described as:

Choose a random layer 0 ≤ i < n.
Let x = U₀x_i.
If x<x_i+1, return x.
Let y = y_i + U₁(y_i+1 − y_i).
Compute f(x). If y<f(x), return x.
Otherwise, choose new random numbers and go back to step 1.

Step 1 amounts to choosing a low-resolution y coordinate. Step 3 tests if the x coordinate is clearly within the desired density function without knowing more about the y coordinate. If it is not, step 4 chooses a high-resolution y coordinate, and step 5 does the rejection test.

With closely spaced layers, the algorithm terminates at step 3 a very large fraction of the time. For the top layer n − 1, however, this test always fails, because x_n = 0.

Layer 0 can also be divided into a central region and an edge, but the edge is an infinite tail. To use the same algorithm to check if the point is in the central region, generate a fictitious x₀ = A/y₁. This will generate points with x < x₁ with the correct frequency, and in the rare case that layer 0 is selected and x ≥ x₁, use a special fallback algorithm to select a point at random from the tail. Because the fallback algorithm is used less than one time in a thousand, speed is not essential.

Thus, the full ziggurat algorithm for one-sided distributions is:

Choose a random layer 0 ≤ i < n.
Let x = U₀x_i
If x<x_i+1, return x.
If i = 0, generate a point from the tail using the fallback algorithm.
Let y = y_i + U₁(y_i+1 − y_i).
Compute f(x). If y<f(x), return x.
Otherwise, choose new random numbers and go back to step 1.

For a two-sided distribution, the result must be negated 50% of the time. This can often be done conveniently by choosing U₀ ∈ (−1,1) and, in step 3, testing if |x| <x_i+1.

Fallback algorithms for the tail

Because the ziggurat algorithm only generates most outputs very rapidly, and requires a fallback algorithm whenever x > x₁, it is always more complex than a more direct implementation. The specific fallback algorithm depends on the distribution.

For an exponential distribution, the tail looks just like the body of the distribution. One way is to fall back to the most elementary algorithm E = −ln(U₁) and let x = x₁ − ln(U₁). Another is to call the ziggurat algorithm recursively and add x₁ to the result.

For a normal distribution, Marsaglia suggests a compact algorithm:

Let x = −ln(U₁)/x₁.
Let y = −ln(U₂).
If 2y > x², return x + x₁.
Otherwise, go back to step 1.

Since x₁ ≈ 3.5 for typical table sizes, the test in step 3 is almost always successful.

Optimizations

The algorithm can be performed efficiently with precomputed tables of x_i and y_i = f(x_i), but there are some modifications to make it even faster:

Nothing in the ziggurat algorithm depends on the probability distribution function being normalized (integral under the curve equal to 1), removing normalizing constants can speed up the computation of f(x).
Most uniform random number generators are based on integer random number generators which return an integer in the range [0, 2³² − 1]. A table of 2⁻³²x_i lets you use such numbers directly for U₀.
When computing two-sided distributions using a two-sided U₀ as described earlier, the random integer can be interpreted as a signed number in the range [−2³¹, 2³¹ − 1], and a scale factor of 2⁻³¹ can be used.
Rather than comparing U₀x_i to x_i+1 in step 3, it is possible to precompute x_i+1/x_i and compare U₀ with that directly. If U₀ is an integer random number generator, these limits may be premultiplied by 2³² (or 2³¹, as appropriate) so an integer comparison can be used.
With the above two changes, the table of unmodified x_i values is no longer needed and may be deleted.
When generating IEEE 754 single-precision floating point values, which only have a 24-bit mantissa (including the implicit leading 1), the least-significant bits of a 32-bit integer random number are not used. These bits may be used to select the layer number. (See the references below for a detailed discussion of this.)
The first three steps may be put into an inline function, which can call an out-of-line implementation of the less frequently needed steps.

Generating the tables

It is possible to store the entire table precomputed, or just include the values n, y₁, A, and an implementation of f⁻¹(y) in the source code, and compute the remaining values when initializing the random number generator.

As previously described, you can find x_i = f⁻¹(y_i) and y_i+1 = y_i + A/x_i. Repeat n − 1 times for the layers of the ziggurat. At the end, you should have y_n = f(0). There will be some round-off error, but it is a useful sanity test to see that it is acceptably small.

When actually filling in the table values, just assume that x_n = 0 and y_n = f(0), and accept the slight difference in layer n − 1's area as rounding error.

Finding x₁ and A

Given an initial (guess at) x₁, you need a way to compute the area t of the tail for which x > x₁. For the exponential distribution, this is just e^−x₁, while for the normal distribution, assuming you are using the unnormalized f(x) = e^−x²/2, this is √π/2 erfc(x/√2). For more awkward distributions, numerical integration may be required.

With this in hand, from x₁, you can find y₁ = f(x₁), the area t in the tail, and the area of the base layer A = x₁y₁ + t.

Then compute the series y_i and x_i as above. If y_i > f(0) for any i<n, then the initial estimate x₁ was too low, leading to too large an area A. If y_n<f(0), then the initial estimate x₁ was too high.

Given this, use a root-finding algorithm (such as the bisection method) to find the value x₁ which produces y_n−1 as close to f(0) as possible. Alternatively, look for the value which makes the area of the topmost layer, x_n−1(f(0) − y_n−1), as close to the desired value A as possible. This saves one evaluation of f⁻¹(x) and is actually the condition of greatest interest.

Related Research Articles

In mathematics and computer programming, exponentiating by squaring is a general method for fast computation of large positive integer powers of a number, or more generally of an element of a semigroup, like a polynomial or a square matrix. Some variants are commonly referred to as square-and-multiply algorithms or binary exponentiation. These can be of quite general use, for example in modular arithmetic or powering of matrices. For semigroups for which additive notation is commonly used, like elliptic curves used in cryptography, this method is also referred to as double-and-add.

In probability theory, a probability density function (PDF), density function, or density of an absolutely continuous random variable, is a function whose value at any given sample in the sample space can be interpreted as providing a relative likelihood that the value of the random variable would be equal to that sample. Probability density is the probability per unit length, in other words, while the absolute likelihood for a continuous random variable to take on any particular value is 0, the value of the PDF at two different samples can be used to infer, in any particular draw of the random variable, how much more likely it is that the random variable would be close to one sample compared to the other sample.

A pseudorandom number generator (PRNG), also known as a deterministic random bit generator (DRBG), is an algorithm for generating a sequence of numbers whose properties approximate the properties of sequences of random numbers. The PRNG-generated sequence is not truly random, because it is completely determined by an initial value, called the PRNG's seed. Although sequences that are closer to truly random can be generated using hardware random number generators, pseudorandom number generators are important in practice for their speed in number generation and their reproducibility.

The Box–Muller transform, by George Edward Pelham Box and Mervin Edgar Muller, is a random number sampling method for generating pairs of independent, standard, normally distributed random numbers, given a source of uniformly distributed random numbers. The method was in fact first mentioned explicitly by Raymond E. A. C. Paley and Norbert Wiener in 1934.

Bresenham's line algorithm is a line drawing algorithm that determines the points of an n-dimensional raster that should be selected in order to form a close approximation to a straight line between two points. It is commonly used to draw line primitives in a bitmap image, as it uses only integer addition, subtraction, and bit shifting, all of which are very cheap operations in historically common computer architectures. It is an incremental error algorithm, and one of the earliest algorithms developed in the field of computer graphics. An extension to the original algorithm may be used for drawing circles.

In mathematics, a Diophantine equation is an equation of the form P(x₁, ..., x_j, y₁, ..., y_k) = 0 (usually abbreviated P(x, y) = 0) where P(x, y) is a polynomial with integer coefficients, where x₁, ..., x_j indicate parameters and y₁, ..., y_k indicate unknowns.

In computer graphics, a line drawing algorithm is an algorithm for approximating a line segment on discrete graphical media, such as pixel-based displays and printers. On such media, line drawing requires an approximation. Basic algorithms rasterize lines in one color. A better representation with multiple color gradations requires an advanced process, spatial anti-aliasing.

Perlin noise is a type of gradient noise developed by Ken Perlin in 1983. It has many uses, including but not limited to: procedurally generating terrain, applying pseudo-random changes to a variable, and assisting in the creation of image textures. It is most commonly implemented in two, three, or four dimensions, but can be defined for any number of dimensions.

In probability theory, coupling is a proof technique that allows one to compare two unrelated random variables (distributions) $X$ and $Y$ by creating a random vector $W$ whose marginal distributions correspond to $X$ and $Y$ respectively. The choice of $W$ is generally not unique, and the whole idea of "coupling" is about making such a choice so that $X$ and $Y$ can be related in a particularly desirable way.

A random permutation is a random ordering of a set of objects, that is, a permutation-valued random variable. The use of random permutations is often fundamental to fields that use randomized algorithms such as coding theory, cryptography, and simulation. A good example of a random permutation is the shuffling of a deck of cards: this is ideally a random permutation of the 52 cards.

George Marsaglia was an American mathematician and computer scientist. He is best known for creating the diehard tests, a suite of software for measuring statistical randomness.

Random number generation is a process by which, often by means of a random number generator (RNG), a sequence of numbers or symbols that cannot be reasonably predicted better than by random chance is generated. This means that the particular outcome sequence will contain some patterns detectable in hindsight but unpredictable to foresight. True random number generators can be hardware random-number generators (HRNGs), wherein each generation is a function of the current value of a physical environment's attribute that is constantly changing in a manner that is practically impossible to model. This would be in contrast to so-called "random number generations" done by pseudorandom number generators (PRNGs), which generate numbers that only look random but are in fact pre-determined—these generations can be reproduced simply by knowing the state of the PRNG.

Random self-reducibility (RSR) is the rule that a good algorithm for the average case implies a good algorithm for the worst case. RSR is the ability to solve all instances of a problem by solving a large fraction of the instances.

In computer science, multiply-with-carry (MWC) is a method invented by George Marsaglia for generating sequences of random integers based on an initial set from two to many thousands of randomly chosen seed values. The main advantages of the MWC method are that it invokes simple computer integer arithmetic and leads to very fast generation of sequences of random numbers with immense periods, ranging from around $to .$

<span class="mw-page-title-main">Truncated normal distribution</span> Type of probability distribution

In probability and statistics, the truncated normal distribution is the probability distribution derived from that of a normally distributed random variable by bounding the random variable from either below or above. The truncated normal distribution has wide applications in statistics and econometrics.

In computer graphics, a digital differential analyzer (DDA) is hardware or software used for interpolation of variables over an interval between start and end point. DDAs are used for rasterization of lines, triangles and polygons. They can be extended to non linear functions, such as perspective correct texture mapping, quadratic curves, and traversing voxels.

The Lehmer random number generator, sometimes also referred to as the Park–Miller random number generator, is a type of linear congruential generator (LCG) that operates in multiplicative group of integers modulo n. The general formula is

Non-uniform random variate generation or pseudo-random number sampling is the numerical practice of generating pseudo-random numbers (PRN) that follow a given probability distribution. Methods are typically based on the availability of a uniformly distributed PRN generator. Computational algorithms are then used to manipulate a single random variate, X, or often several such variates, into a new random variate Y such that these values have the required distribution. The first methods were developed for Monte-Carlo simulations in the Manhattan project, published by John von Neumann in the early 1950s.

In computing, the alias method is a family of efficient algorithms for sampling from a discrete probability distribution, published in 1974 by A. J. Walker. That is, it returns integer values $1 \leq i \leq n$ according to some arbitrary probability distribution $p i$ . The algorithms typically use $O (n log n)$ or $O (n)$ preprocessing time, after which random values can be drawn from the distribution in $O (1)$ time.

Digital signatures are a means to protect digital information from intentional modification and to authenticate the source of digital information. Public key cryptography provides a rich set of different cryptographic algorithms the create digital signatures. However, the primary public key signatures currently in use will become completely insecure if scientists are ever able to build a moderately sized quantum computer. Post quantum cryptography is a class of cryptographic algorithms designed to be resistant to attack by a quantum cryptography. Several post quantum digital signature algorithms based on hard problems in lattices are being created replace the commonly used RSA and elliptic curve signatures. A subset of these lattice based scheme are based on a problem known as Ring learning with errors. Ring learning with errors based digital signatures are among the post quantum signatures with the smallest public key and signature sizes

References

George Marsaglia; Wai Wan Tsang (2000). "The Ziggurat Method for Generating Random Variables". Journal of Statistical Software. 5 (8). Retrieved 2007-06-20. This paper numbers the layers from 1 starting at the top, and makes layer 0 at the bottom a special case, while the explanation above numbers layers from 0 at the bottom.
C implementation of the ziggurat method for the normal density function and the exponential density function, that is essentially a copy of the code in the paper. (Potential users should be aware that this C code assumes 32-bit integers.)
A C# implementation of the ziggurat algorithm and overview of the method.
Jurgen A. Doornik (2005). "An Improved Ziggurat Method to Generate Normal Random Samples" (PDF). Nuffield College, Oxford. Retrieved 2007-06-20. Describes the hazards of using the least-significant bits of the integer random number generator to choose the layer number.
Normal Behavior By Cleve Moler, MathWorks, describing the ziggurat algorithm introduced in MATLAB version 5, 2001.
The Ziggurat Random Normal Generator Blogs of MathWorks, posted by Cleve Moler, May 18, 2015.
David B. Thomas; Philip H.W. Leong; Wayne Luk; John D. Villasenor (October 2007). "Gaussian Random Number Generators" (PDF). ACM Computing Surveys. 39 (4): 11:1–38. doi:10.1145/1287620.1287622. ISSN 0360-0300. S2CID 10948255 . Retrieved 2009-07-27. [W]hen maintaining extremely high statistical quality is the first priority, and subject to that constraint, speed is also desired, the Ziggurat method will often be the most appropriate choice. Comparison of several algorithms for generating Gaussian random numbers.
Nadler, Boaz (2006). "Design Flaws in the Implementation of the Ziggurat and Monty Python methods (And some remarks on Matlab randn)". arXiv: math/0603058 .. Illustrates problems with underlying uniform pseudo-random number generators and how those problems affect the ziggurat algorithm's output.
Edrees, Hassan M.; Cheung, Brian; Sandora, McCullen; Nummey, David; Stefan, Deian (13–16 July 2009). Hardware-Optimized Ziggurat Algorithm for High-Speed Gaussian Random Number Generators (PDF). 2009 International Conference on Engineering of Reconfigurable Systems & Algorithms. Las Vegas.
Marsaglia, George (September 1963). Generating a Variable from the Tail of the Normal Distribution (Technical report). Boeing Scientific Research Labs. Mathematical Note No. 322, DTIC accession number AD0423993. Archived from the original on September 10, 2014 – via Defense Technical Information Center.

This page is based on this Wikipedia article
Text is available under the CC BY-SA 4.0 license; additional terms may apply.
Images, videos and audio are available under their respective licenses.