10.1 Facts About the Chi-Square Distribution

Aligned outcomes:

SLO 4

Demonstrate an understanding of, and ability to use, basic ideas of statistical processes, including hypothesis tests and confidence interval estimation.

SLO 5

Identify appropriate statistical techniques and use technology-based statistical analysis to describe, interpret, and communicate results.

Learning Objectives

By the end of this section, you will be able to:

In this section, you will learn to:
  • write the notation for a chi-square distribution and say what its degrees of freedom stand for;
  • compute the mean and the standard deviation of a chi-square distribution from its degrees of freedom;
  • describe a chi-square random variable as a sum of squared independent standard normal variables;
  • list the properties of the chi-square curve, including its skew, its floor at zero, and the normal approximation that takes over when the degrees of freedom get large.

Every distribution you have met so far was introduced because some question needed it. The normal curve answered questions about a single measurement; the Student's t curve answered them again when the standard deviation had to be estimated. The chi-square distribution answers a different kind of question altogether — questions about counts that fall into categories, and questions about how spread out a population is. Before any of those tests will make sense, you need to know what the curve itself looks like and where its numbers come from. That is all this section does. There is no test to run yet, only a distribution to get acquainted with, and the acquaintance is short: one piece of notation, two formulas, and five facts about the shape.

10.1.1 The Notation and Its Degrees of Freedom

The chi-square distribution is written with the Greek letter chi, \(\chi\), carrying a superscript 2. The 2 is part of the name, not an instruction to square anything — you will see shortly that the squaring has already happened by the time the variable exists.

Definition 10.1.1: Chi-Square Distribution

A random variable \(X\) follows a chi-square distribution with \(df\) degrees of freedom when we write

$$ X \sim \chi_{df}^{2} $$

The subscript \(df\) is the one number that decides the shape of the curve.

The random variable is shown here as \(\chi^{2}\), but it may be written with any upper-case letter. Seeing \(X\) or \(Y\) on one page and \(\chi^{2}\) on the next does not signal a different distribution; it is only a change of lettering, the same way \(x\) and \(t\) both name a variable in algebra.

Definition 10.1.1 — Chi-Square Distribution The notation X ~ chi-squared subscript df appears at the top. Below, a single horizontal axis carries three chi-square density curves drawn to scale for df = 1, df = 4, and df = 10. The df = 1 curve plunges from the top of the frame near zero and dies out quickly; the df = 4 curve peaks early and skews right; the df = 10 curve peaks farther right and is nearly symmetric. Each curve is labelled with its chi-squared symbol and its df value. A caption states that the subscript df is the one number that decides the shape of the curve. X ~ χ²df one number, df, chooses the curve x 2 4 6 8 10 12 14 χ²1, df = 1 χ²4, df = 4 χ²10, df = 10 The subscript df is the one number that decides the shape of the curve.

Definition 10.1.1 — Chi-Square Distribution: The subscript df is the one number that decides the shape of the curve.

Definition 10.1.2: Degrees of Freedom

The degrees of freedom, abbreviated \(df\), is the parameter that fixes which chi-square curve you are working with. Its value depends on how the chi-square distribution is being used.

Why a parameter deserves this much attention

With the normal distribution you had to supply two numbers, a mean and a standard deviation, before the curve was pinned down. Chi-square asks for only one. That single number does all the work — it sets the center, the spread, and the amount of skew at the same time. So when a problem hands you a chi-square distribution, finding \(df\) is not a preliminary step before the real work; it is most of the work.

If you simply want to practice calculating chi-square probabilities, use \(df = n - 1\). Be careful not to read that as a universal rule. The three major uses of the chi-square distribution each calculate degrees of freedom differently, and you will meet those rules one at a time in the sections that follow. What matters right now is the idea: \(df\) is not decoration attached to the curve, it is the curve. Change it and you are looking at a different distribution.

Definition 10.1.2 — Degrees of freedom The horizontal axis shows values of X from 0 to 30. Three chi-square density curves are drawn on the same axes: chi-square with 1 degree of freedom rises steeply at zero and falls fast; chi-square with 4 degrees of freedom peaks at x = 2; chi-square with 10 degrees of freedom peaks at x = 8 with a longer right tail. Each curve carries its own notation label, showing that changing the single parameter df changes which curve you are looking at. 0 5 10 15 20 25 30 x χ²₁ χ²₄ χ²₁₀ One number — df — selects the whole curve.

Definition 10.1.2 — Degrees of Freedom: The horizontal axis shows values of X from 0 to 30.

Try It Now 10.1.1

Write the notation for a chi-square distribution with 12 degrees of freedom, and say what the 12 refers to.

Solution — reading and writing the notation

Step 1 — Start with the symbol for the distribution. The chi-square distribution is written \(\chi^{2}\).

Step 2 — Attach the degrees of freedom as a subscript. The degrees of freedom go underneath, giving \(\chi_{12}^{2}\).

Step 3 — Name the random variable. Following the pattern in Definition 10.1.1, a variable \(X\) drawn from this distribution is written

$$ X \sim \chi_{12}^{2} $$

Step 4 — Say what the 12 means. The 12 is the degrees of freedom. It selects one particular chi-square curve out of the whole family of them.

Answer: \(X \sim \chi_{12}^{2}\); the 12 is the degrees of freedom, and it identifies which chi-square curve is being used.

Example 10.1.1: Two variables, one distribution

A classmate writes \(Y \sim \chi_{7}^{2}\) while the textbook writes \(\chi_{7}^{2}\) with no letter in front of it at all. Are these describing the same thing?

Solution — the letter is not the distribution

Step 1 — Compare the distributions. Both carry the symbol \(\chi^{2}\) with a subscript of 7, so both name the chi-square distribution with 7 degrees of freedom.

Step 2 — Compare the variable names. Your classmate has given the random variable the name \(Y\). The textbook has not bothered to name it. The random variable may be shown as \(\chi^{2}\) itself, or as any upper-case letter.

Step 3 — Decide. Nothing that distinguishes one distribution from another has changed. Only the labeling has.

Answer: Yes. Both describe a chi-square distribution with \(df = 7\); the choice of letter carries no information.

10.1.2 The Mean and the Standard Deviation

Because a single parameter controls the whole curve, the center and the spread both fall out of that one number.

For a \(\chi^{2}\) distribution, the population mean is

$$ \mu = df $$

and the population standard deviation is

$$ \sigma = \sqrt{2(df)} $$

The mean is the easier of the two to remember, because there is nothing to remember: the mean is the degrees of freedom. A chi-square distribution with 30 degrees of freedom has a mean of 30.

The standard deviation grows too, but much more slowly. Doubling the degrees of freedom doubles the mean, while it multiplies the standard deviation by only \(\sqrt{2} \approx 1.41\). That mismatch is worth holding on to: as \(df\) climbs, the curve slides to the right faster than it spreads out, which makes it look proportionally tighter around its center even as the raw spread increases.

The units are not the same on both sides

It is tempting to compare \(\mu = df\) and \(\sigma = \sqrt{2(df)}\) directly and conclude that the standard deviation is always "smaller." For \(df = 2\), though, \(\mu = 2\) and \(\sigma = 2\) exactly, and for \(df = 1\) the standard deviation \(\sqrt{2} \approx 1.41\) is larger than the mean. The standard deviation only becomes small relative to the mean once \(df\) is past 2, and it keeps shrinking in relative terms from there.

Try It Now 10.1.2

A chi-square distribution has \(df = 15\). Find its mean and its standard deviation, rounding the standard deviation to two decimal places.

Solution — mean and standard deviation from df alone

Step 1 — Write both formulas down.

$$ \mu = df \qquad \text{and} \qquad \sigma = \sqrt{2(df)} $$

Step 2 — Substitute \(df = 15\) into the mean. The mean is the degrees of freedom, so \(\mu = 15\).

Step 3 — Substitute \(df = 15\) into the standard deviation.

$$ \sigma = \sqrt{2(15)} = \sqrt{30} \approx 5.4772 $$

Step 4 — Round as asked. To two decimal places, \(\sigma \approx 5.48\).

Answer: \(\mu = 15\) and \(\sigma \approx 5.48\).

Example 10.1.2: Comparing two curves

Two chi-square distributions have \(df = 8\) and \(df = 32\). Find the mean and standard deviation of each, then describe how the second curve differs from the first.

Solution — what quadrupling the degrees of freedom does

Step 1 — Handle the first curve. With \(df = 8\),

$$ \mu = 8 \qquad \sigma = \sqrt{2(8)} = \sqrt{16} = 4 $$

Step 2 — Handle the second curve. With \(df = 32\),

$$ \mu = 32 \qquad \sigma = \sqrt{2(32)} = \sqrt{64} = 8 $$

Step 3 — Compare. The degrees of freedom were multiplied by 4. The mean was also multiplied by 4, moving from 8 to 32. The standard deviation was multiplied by only 2, moving from 4 to 8.

Step 4 — Put it in words. The second curve sits much further to the right, and although it is genuinely wider, it is narrower in proportion to its own mean. That is the visual signature of a chi-square curve becoming less skewed as \(df\) rises.

Answer: \(df = 8\) gives \(\mu = 8\), \(\sigma = 4\); \(df = 32\) gives \(\mu = 32\), \(\sigma = 8\). Quadrupling \(df\) quadruples the mean but only doubles the standard deviation.

10.1.3 What a Chi-Square Variable Is Built From

So far the chi-square distribution has been described from the outside — a symbol, a parameter, two formulas. The definition below says what it actually is, and every property in the next section follows from it.

Definition 10.1.3: Chi-Square Random Variable

The random variable for a chi-square distribution with \(k\) degrees of freedom is the sum of \(k\) independent, squared standard normal variables:

$$ \chi^{2} = (Z_{1})^{2} + (Z_{2})^{2} + \dots + (Z_{k})^{2} $$

where each \(Z_{i}\) is an independent standard normal variable.

Where the word "freedom" comes from

Degrees of freedom counts the independent pieces of information going into the statistic. In the construction above the counting is literal: each independent standard normal contributes exactly one degree of freedom, so summing \(k\) of them produces a distribution with \(k\) degrees of freedom. When you meet \(df = n - 1\) later, the subtraction is there because one piece of information has already been spent estimating something else.

Read that construction slowly, because three of the five facts in the next section are already hiding inside it.

Each \(Z\) is a standard normal variable — the familiar bell curve, centered at zero, running off to negative values as readily as positive ones. Squaring destroys the sign. A \(Z\) of \(-2.1\) and a \(Z\) of \(+2.1\) both contribute \(4.41\) to the sum, and no single term can ever contribute a negative amount. Add \(k\) such terms and the total cannot be negative either.

Squaring also stretches the large values much harder than the small ones. A \(Z\) of \(0.5\) shrinks to \(0.25\), while a \(Z\) of \(3\) grows to \(9\). Values near zero get pulled toward zero and unusual values get pushed far out to the right, which is exactly how a long right tail is manufactured.

Definition 10.1.3 — Chi-Square Random Variable Five standard normal variables Z1 through Z5 with sample values -2.1, 0.5, 3, -1.2 and 1.8 are each squared to give 4.41, 0.25, 9, 1.44 and 3.24. The squares are added to give chi-squared equal to 18.34, a chi-square distribution with 5 degrees of freedom. standard normal squared — sign destroyed sum of k terms Z₁ = −2.1 4.41 Z₂ = 0.5 0.25 Z₃ = 3 9 Z₄ = −1.2 1.44 Z₅ = 1.8 3.24 + + + + χ² = 18.34 k = 5 degrees of freedom Every term is a square, so every term ≥ 0 — and so is the sum. Squaring stretches large |Z| far rightward: 3 becomes 9, while 0.5 shrinks to 0.25. Each independent Z contributes one degree of freedom.

Definition 10.1.3 — Chi-Square Random Variable: Every term is a square, so every term ≥ 0 — and so is the sum. Squaring stretches large |Z| far rightward: 3 becomes 9, while 0.5 shrinks to 0.25.....

Try It Now 10.1.3

Suppose \(Z_{1}, Z_{2}, Z_{3}, Z_{4}\) and \(Z_{5}\) are independent standard normal variables. Name the distribution of \(Z_{1}^{2} + Z_{2}^{2} + Z_{3}^{2} + Z_{4}^{2} + Z_{5}^{2}\), and give its mean.

Solution — counting the terms

Step 1 — Check the construction matches the definition. Each term is an independent standard normal that has been squared, and the terms are added. That is exactly the form in Definition 10.1.3.

Step 2 — Count the terms. There are five squared standard normals, so \(k = 5\).

Step 3 — Name the distribution. The sum follows a chi-square distribution with 5 degrees of freedom, written \(\chi_{5}^{2}\).

Step 4 — Find the mean. The mean of a chi-square distribution is its degrees of freedom, so \(\mu = 5\).

Answer: The sum follows \(\chi_{5}^{2}\), with mean \(\mu = 5\).

Example 10.1.3: Why the total can never come out negative

A student calculates a chi-square test statistic and gets \(-3.2\). Without knowing anything else about the problem, what can you conclude?

Solution — using the construction as a check

Step 1 — Recall how the statistic is built. By Definition 10.1.3, a chi-square variable is a sum of squared quantities.

Step 2 — Ask what a squared quantity can be. Squaring any real number gives a result that is zero or positive. A negative result is impossible.

Step 3 — Ask what a sum of such quantities can be. Adding terms that are each zero or positive gives a total that is zero or positive.

Step 4 — Conclude. A chi-square value of \(-3.2\) cannot occur. The student has made an arithmetic error somewhere — most often a subtraction carried out in the wrong order, or a squaring step that was skipped.

Answer: The result is impossible; a chi-square statistic is always greater than or equal to zero, so there is a computational mistake to find.

10.1.4 Five Facts About the Curve

These five facts describe every chi-square curve you will meet. Each one traces back to the construction in Definition 10.1.3.

1. The curve is nonsymmetrical and skewed to the right. It is not the mirror-image bell you are used to. It begins at zero, rises to a peak, then trails away slowly to the right. The squaring step is responsible: it pins the left side at zero while stretching unusual values far out to the right.

2. There is a different chi-square curve for each df. The degrees of freedom do not adjust a fixed shape, they select a different one. A curve with small \(df\) is severely skewed and peaks close to zero; as \(df\) grows the peak slides right and the skew fades.

3. The test statistic for any test is always greater than or equal to zero. This follows directly from Example 10.1.3, and it makes a free error check available on every calculation you do in this chapter.

4. When \(df > 90\), the chi-square curve approximates the normal distribution. The next section works through what that means and when to lean on it.

5. The mean, \(\mu\), is located just to the right of the peak. In any right-skewed distribution the long tail drags the mean away from the peak, so the mean sits at a slightly larger value than the mode.

Table 10.1.1 — The five facts, and where each one comes from.
FactWhy it holds
Nonsymmetrical, skewed rightSquaring pins the left side at zero and stretches large values rightward.
A different curve for each \(df\)\(df\) counts the terms in the sum, so it changes the variable itself.
Always \(\geq 0\)A sum of squares cannot be negative.
Approximately normal when \(df > 90\)A sum of many independent terms tends toward a normal shape.
Mean just right of the peakThe long right tail pulls the mean above the mode.
A sanity check you get for free

Fact 3 is the most immediately useful of the five. Chi-square calculations involve a lot of subtracting, squaring and dividing, and a sign slip is easy to make and hard to spot. Any negative total is proof of an error before you have looked at a single table value.

Try It Now 10.1.4

For a chi-square distribution with \(df = 8\), state the mean and the standard deviation, describe the shape of the curve, and say where the mean sits relative to the peak.

Solution — reading four facts off one parameter

Step 1 — Find the mean. The mean equals the degrees of freedom, so \(\mu = 8\).

Step 2 — Find the standard deviation.

$$ \sigma = \sqrt{2(8)} = \sqrt{16} = 4 $$

Step 3 — Describe the shape. Every chi-square curve is nonsymmetrical and skewed to the right, and \(df = 8\) is small enough that the skew is pronounced.

Step 4 — Locate the mean against the peak. Because the curve is skewed right, the mean lies just to the right of the peak.

Answer: \(\mu = 8\), \(\sigma = 4\); the curve is skewed to the right, with the mean just right of the peak.

Example 10.1.4: Judging a statistic without a table

A chi-square test produces a statistic of \(28.7\) with \(df = 14\). Using only the mean and standard deviation, decide whether that value looks ordinary or unusually large.

Solution — measuring a statistic in standard deviations

Step 1 — Find the center of the distribution. With \(df = 14\), the mean is \(\mu = 14\).

Step 2 — Find the spread.

$$ \sigma = \sqrt{2(14)} = \sqrt{28} \approx 5.29 $$

Step 3 — Measure the distance in standard deviations.

$$ \frac{28.7 - 14}{5.29} \approx 2.78 $$

Step 4 — Interpret. The statistic sits about 2.8 standard deviations above the mean. The chi-square curve is skewed right, so its right tail is thinner than a normal curve's at the same distance, which makes a value this far out uncommon.

Answer: \(28.7\) is unusually large for \(df = 14\) — roughly 2.8 standard deviations above a mean of 14.

10.1.5 When the Degrees of Freedom Get Large

Fact 4 said the chi-square curve starts to look normal once \(df\) passes about 90. That is more than a visual resemblance; it licenses a genuine shortcut.

When \(df > 90\), a chi-square distribution is approximately normal, with the same mean and standard deviation the chi-square formulas already give you:

$$ \mu = df \qquad \sigma = \sqrt{2(df)} $$

Take \(X \sim \chi_{1{,}000}^{2}\). The mean is \(\mu = df = 1{,}000\) and the standard deviation is \(\sigma = \sqrt{2(1{,}000)} = \sqrt{2{,}000} \approx 44.7\). Therefore \(X \sim N(1{,}000, 44.7)\), approximately.

Why the skew fades

A chi-square variable with \(k\) degrees of freedom is a sum of \(k\) independent pieces. Sums of many independent random quantities drift toward a normal shape no matter how lopsided the individual pieces are — the same behavior that drove the Central Limit Theorem. At \(df = 2\) the sum has two terms and is violently skewed; at \(df = 1{,}000\) it has a thousand, and the skew has all but disappeared. The threshold of 90 is a practical convention, not a point where anything abruptly changes.

Try It Now 10.1.5

A chi-square distribution has \(df = 100\). Explain whether the normal approximation applies, and if it does, state the approximating normal distribution.

Solution — checking the condition, then naming the curve

Step 1 — Check the condition. The approximation applies when \(df > 90\). Here \(df = 100\), which is greater than 90, so it applies.

Step 2 — Find the mean. \(\mu = df = 100\).

Step 3 — Find the standard deviation.

$$ \sigma = \sqrt{2(100)} = \sqrt{200} \approx 14.14 $$

Step 4 — State the approximating distribution. \(X \sim N(100, 14.14)\), approximately.

Answer: The approximation applies because \(df = 100 > 90\), and \(X \sim N(100, 14.14)\).

Example 10.1.5: A distribution below the threshold

A chi-square distribution has \(df = 40\). Does the normal approximation apply? Describe the shape of this curve and compare it with the curve for \(df = 4\).

Solution — "less skewed" is not "normal"

Step 1 — Check the condition. The rule requires \(df > 90\). Here \(df = 40\), so the normal approximation does not apply.

Step 2 — Describe the shape anyway. The curve is still nonsymmetrical and skewed to the right, as every chi-square curve is.

Step 3 — Compare with \(df = 4\). Both are skewed right, but the skew softens as \(df\) rises, so the \(df = 40\) curve is noticeably closer to symmetric than the \(df = 4\) curve.

Step 4 — State the distinction carefully. Being closer to symmetric is not the same as being normal. The \(df = 40\) curve is nearer a normal shape than the \(df = 4\) curve, but not near enough to substitute one for the other.

Answer: The approximation does not apply at \(df = 40\). The curve is skewed right, less sharply than at \(df = 4\), but still not close enough to normal to be treated as normal.

Problem Set 10.1

Problem 1. Write the notation for a chi-square distribution with 8 degrees of freedom.

Solution

Step 1 — Start with the chi-square symbol: The distribution is written \(\chi^{2}\).

Step 2 — Attach the degrees of freedom as a subscript: With \(df = 8\), the notation becomes \(\chi_{8}^{2}\).

Step 3 — Name the random variable: Following the pattern \(X \sim \chi_{df}^{2}\), a variable from this distribution is written

$$ X \sim \chi_{8}^{2} $$

Answer: \(X \sim \chi_{8}^{2}\), the chi-square distribution with 8 degrees of freedom.

Problem 2. Find the mean and the standard deviation of a chi-square distribution with \(df = 20\). Round the standard deviation to two decimal places.

Solution

Step 1 — Write both formulas. For any chi-square distribution,

$$ \mu = df \qquad \text{and} \qquad \sigma = \sqrt{2(df)} $$

The mean is simply the degrees of freedom, and the standard deviation comes from doubling the degrees of freedom and taking the square root.

Step 2 — Find the mean. Substituting \(df = 20\),

$$ \mu = 20 $$

Step 3 — Find the standard deviation.

$$ \sigma = \sqrt{2(20)} = \sqrt{40} \approx 6.3246 $$

Step 4 — Round as asked. To two decimal places, \(\sigma \approx 6.32\).

Answer: \(\mu = 20\) and \(\sigma \approx 6.32\).

Problem 3. If \(Z_{1}, Z_{2}, Z_{3}\) and \(Z_{4}\) are independent standard normal variables, what is the distribution of \(Z_{1}^{2} + Z_{2}^{2} + Z_{3}^{2} + Z_{4}^{2}\)?

Solution

Step 1 — Compare with the definition. A chi-square random variable with \(k\) degrees of freedom is the sum of \(k\) independent, squared standard normal variables: \(\chi^{2} = (Z_{1})^{2} + \dots + (Z_{k})^{2}\). The given expression has exactly this form.

Step 2 — Count the terms. There are four independent squared standard normals in the sum, so \(k = 4\).

Step 3 — Name the distribution. The sum follows a chi-square distribution with 4 degrees of freedom:

$$ Z_{1}^{2} + Z_{2}^{2} + Z_{3}^{2} + Z_{4}^{2} \sim \chi_{4}^{2} $$

(Its mean would be \(\mu = df = 4\).)

Answer: The sum follows a chi-square distribution with 4 degrees of freedom, written \(\chi_{4}^{2}\).

Problem 4. Name two properties of the chi-square curve that distinguish it from the normal curve.

Solution

Step 1 — Recall how each curve is shaped. The normal curve is symmetric about its mean; the chi-square curve is not.

Step 2 — State two distinguishing properties. Any two of the following distinguish the chi-square curve from the normal curve:

  1. Skew: The chi-square curve is nonsymmetrical and skewed to the right, while the normal curve is symmetric.
  2. Floor at zero: Every chi-square value is greater than or equal to zero (it is a sum of squares), whereas a normal variable can take any real value, positive or negative.

(Other acceptable answers: there is a different chi-square curve for each \(df\); when \(df > 90\) the chi-square curve approximates the normal; the mean sits just to the right of the peak.)

Answer: Two valid properties are that the chi-square curve is skewed to the right (not symmetric) and that it never takes negative values, since it is built from sums of squares.

Problem 5. A chi-square distribution has \(df = 120\). Would you expect the normal approximation to be good? If so, state the mean and the standard deviation of the approximating normal distribution.

Solution

Step 1 — Check the condition for the approximation. The normal approximation applies when \(df > 90\). Since \(120 > 90\), we would expect the approximation to be good.

Step 2 — Find the mean of the approximating normal distribution.

$$ \mu = df = 120 $$

Step 3 — Find the standard deviation of the approximating normal distribution.

$$ \sigma = \sqrt{2(df)} = \sqrt{2(120)} = \sqrt{240} \approx 15.49 $$

Answer: Yes, the normal approximation should be good because \(df = 120 > 90\); the approximating distribution is approximately \(N(120, 15.49)\).

Problem 6. A classmate reports a chi-square test statistic of \(-1.4\). Explain how you know this is wrong without seeing any of the calculations.

Solution

Step 1 — Recall what a chi-square statistic is made of. By Definition 10.1.3, a chi-square variable is a sum of squared terms: \((Z_{1})^{2} + (Z_{2})^{2} + \dots + (Z_{k})^{2}\).

Step 2 — Ask what each term can be. Squaring any real number gives zero or a positive result — never a negative one.

Step 3 — Ask what the sum can be. Adding quantities that are each zero or positive produces a total that is zero or positive. So every chi-square test statistic must satisfy \(\chi^{2} \geq 0\).

Step 4 — Conclude. Since \(-1.4 < 0\), this value cannot be a legitimate chi-square statistic. Your classmate must have made an arithmetic error — most likely a subtraction performed in the wrong order or a squaring step skipped.

Answer: A chi-square statistic is a sum of squares and therefore can never be negative, so \(-1.4\) is impossible regardless of the calculations.

Problem 7. For a chi-square distribution with \(df = 50\), find the mean and the standard deviation. Then say whether the mean is to the left or the right of the peak, and why.

Solution

Step 1 — Find the mean. The mean of a chi-square distribution equals its degrees of freedom:

$$ \mu = df = 50 $$

Step 2 — Find the standard deviation.

$$ \sigma = \sqrt{2(df)} = \sqrt{2(50)} = \sqrt{100} = 10 $$

Step 3 — Decide where the mean sits relative to the peak. Fact 5 says the mean lies just to the right of the peak. This happens because the curve is skewed to the right: the long right tail pulls the mean above the mode (the peak).

Answer: \(\mu = 50\) and \(\sigma = 10\); the mean lies just to the right of the peak because the right skew drags the mean toward the tail.

Problem 8. Explain why a chi-square distribution needs only one parameter when a normal distribution needs two.

Solution

Step 1 — Recall what parameters do for each distribution. A parameter pins down which member of a family of curves you are using. For the normal distribution, you need both a mean \(\mu\) (where the curve is centered) and a standard deviation \(\sigma\) (how spread out it is) before the curve is determined.

Step 2 — Recall how the chi-square curve is built. A chi-square variable is the sum of \(k\) independent, squared standard normal variables. The single number \(k\) — the degrees of freedom — counts those pieces.

Step 3 — See why one number does all the work. Because the construction fixes everything else, \(df\) alone determines the center (\(\mu = df\)), the spread (\(\sigma = \sqrt{2(df)}\)), and the amount of skew simultaneously. Change \(df\) and you get a different curve, but no second parameter is ever needed to finish specifying it.

Answer: A chi-square distribution needs only one parameter because its construction as a sum of \(k\) squared standard normals means the single value \(df = k\) determines the center, the spread, and the skew all at once — unlike the normal curve, whose center and spread can vary independently and so require two numbers.

Key Terms

chi-square distribution — a distribution written \(\chi_{df}^{2}\), formed by adding the squares of independent standard normal variables; it is skewed to the right and its shape is set entirely by its degrees of freedom.

degrees of freedom (\(df\)) — the single parameter of a chi-square distribution; it counts the independent squared standard normal variables in the sum, and it is calculated differently for each major use of the distribution.

chi-square random variable — the sum \((Z_{1})^{2} + (Z_{2})^{2} + \dots + (Z_{k})^{2}\) of \(k\) independent, squared standard normal variables.

test statistic — the value computed from sample data and compared against a chi-square curve; for any chi-square test it is always greater than or equal to zero.

normal approximation — the fact that when \(df > 90\), a chi-square distribution is approximately normal with mean \(\mu = df\) and standard deviation \(\sigma = \sqrt{2(df)}\).