Linear transformation of random variable(s), unbiased estimate of mean and var: notes and practice questions
- An estimator is a random variable; an estimate is its numerical value from a sample.
- An estimator is unbiased (HL) if its expected value equals the population parameter: .
- The sample mean () is always an unbiased estimate for the population mean ().
- The standard sample variance () is not an unbiased estimate for population variance () (HL); it systematically underestimates.
- A linear combination of independent normal variables is also normally distributed (HL).
- For a linear transformation of a single random variable (HL):
- Adding a constant () does not affect variance.
- For linear combinations of multiple independent random variables (HL):
- Always add variances, even when subtracting variables.
- Unbiased estimate of population mean (HL): .
- Unbiased estimate of population variance (HL): .
- Expected value of the sample mean (HL): .
- Variance of the sample mean (HL): .
- Expected value of the sample mean squared (HL): .
- When modeling contextual sums (HL), differentiate scaling a single variable (e.g., ) from summing multiple independent variables (e.g., ).
- HL syllabus explicitly covers algebraic expected value/variance formulas and unbiased estimates.
- GDC notation for unbiased standard deviation is often or ; square it for unbiased variance.
- Never subtract variances; .
- "Sample variance" refers to ; adjust with for the unbiased estimate.
How it is examined
The in the variance transformation is the fact students most often get wrong, and it is a one-mark answer. Distinguishing from on the GDC is the practical skill, since the calculator reports both and only one is the unbiased estimate. Proofs of unbiasedness are not examinable, so a question can only ask for the values.
The unbiased estimate of the population variance, .
- Linear transformation of a single random variable.
- The expected value of linear combinations of random variables.
- The variance of linear combinations of independent random variables.
- as an unbiased estimate of .
- is the variance of the random variable . The variance formula will not be required in examinations.
- **Demonstration that and will not be examined**, but may help understanding.
Linking questions
- TOK: mathematics and the world. In the absence of knowing the value of a parameter, will an unbiased estimator always be better than a biased one?
Practice questions
6 questions · 1 medium · 5 hardQuestion 1
MediumPaper 1 · calculator6 marksA materials scientist is developing a new synthetic fiber and claims its average tensile strength is 120 N. To test this claim, she takes a random sample of 25 fiber segments.
Given that the tensile strength of an individual fiber segment is newtons, the scientist found that, from her 25 samples, and .
(a) Find an unbiased estimate for the mean tensile strength () of the fiber.
(b) Use the formula to determine an unbiased estimate for the variance of the tensile strength of the fiber.
(c) Find a 95% confidence interval for . You may assume that all conditions for a confidence interval have been met.
(d) Suggest, with justification, a valid conclusion that the materials scientist could make regarding her claim.
Recall the formula for the sample mean, which is an unbiased estimate for the population mean.
Carefully substitute the given values into the formula for the unbiased sample variance. Remember to use in the denominator.
You will need the sample mean and the unbiased standard deviation (square root of the variance) from parts (a) and (b). For a 95% confidence interval with a small sample size and unknown population standard deviation, use the t-distribution. You can use your GDC to find the critical t-value or the entire confidence interval.
Compare the claimed average tensile strength (120 N) with the 95% confidence interval you calculated in part (c). If the claimed value falls outside the interval, what does that imply?
Question 2
HardPaper 2 · calculator19 marksA game involves a player attempting to hit a target. The number of successful hits, , in a round follows a discrete probability distribution given by:
(a) Find .
(b) Giving your answers as fractions, calculate:
(i)
(c) The variance of is . Let the random variable .
(i) Find .
(ii) Find .
(d) Let the random variable , be the total of four independent values of .
(i) Find .
(ii) Find .
(e) Let the random variable .
(i) Calculate giving the answer as a fraction.
(ii) Hence, determine whether or not the statement is true. You must justify your answer.
To find , sum the probabilities for and .
The expected value is calculated as .
For a linear transformation , the expected value .
For a linear transformation , the variance .
For independent random variables, the expected value of their sum is the sum of their individual expected values: .
For independent random variables, the variance of their sum is the sum of their individual variances: .
To find , you need to calculate the value of for each value and then use the formula .
Compare your calculated from part (e.i) with the value of using from part (b.i).
Question 3
HardPaper 2 · calculator12 marks(a) A logistics company handles two types of packages: small (S) and large (L). The weight of each type of package follows a normal distribution with parameters as shown in this table:
Type of Package
Mean weight (kg)
Standard deviation (kg)
S (Small)
L (Large)
One package of each type is selected at random. Find the probability that the large package weighs less than four times the weight of the small package.
(b) One large package and three small packages are selected at random. Find the probability that the large package weighs more than the total weight of the three small packages.
Let be the weight of a large package and be the weight of a small package. You need to find . This can be rewritten as . Consider the properties of linear combinations of independent normal random variables to find the mean and variance of .
Let be the weight of a large package and be the weights of three independent small packages. You need to find . First, find the mean and variance of the sum of the three small packages. Then, consider the difference between the large package and this sum.
Question 4
HardPaper 2 · calculator16 marksA factory produces specialized electronic components. The total "quality score" of a component, , is a combination of scores from three independent inspection stages:
- Stage 1: Automated visual inspection. The score from this stage has an expectation of and a standard deviation of .
- Stage 2: Manual functional test. A batch of critical functions are tested, and the score is the number of functions that pass. Each function has a probability of passing, independently.
- Stage 3: Environmental stress test. The score from this stage, representing the number of successful stress cycles, follows a Poisson distribution with a mean of .
The overall quality score for a component is given by .
Calculate the expected value and variance of the total quality score .
Given that the distribution of can be approximated by a Normal distribution, find the probability that a randomly selected component has a total quality score between and (inclusive of , exclusive of ).
The factory manager wants to ensure that the mean total quality score of a sample of components is within units of the true mean, with a probability of at least . Find the minimum sample size required.
Recall the formulas for the expectation and variance of Binomial and Poisson distributions. For independent random variables and constants , and . Remember that .
When approximating a discrete distribution with a continuous Normal distribution, remember to apply a continuity correction. For , the continuous approximation would be .
The Central Limit Theorem states that for a sufficiently large sample size , the sample mean is approximately normally distributed with mean and variance . You will need to use the inverse normal function to find the critical Z-value for the given probability.
Question 5
HardPaper 2 · calculator15 marksA quality control manager at a manufacturing plant wants to assess the consistency of a new batch of electronic components. He decides to test ten components, ensuring that five are selected from Production Line A and five from Production Line B. The manager instructs the supervisors of each line to provide the required number of components from their current production.
(a) Name the type of sampling that best describes the method used by the quality control manager.
The weights, in grams, of the ten components selected for the test are:
.
(b) For these ten components, find
(i) the mean weight.
(ii) the standard deviation of the weights.
The target weight for these components is g. The manager is concerned that the components might be consistently underweight. Perform an appropriate test at the significance level to see if the mean weight of the components produced is less than the target weight. It can be assumed that the weights come from a normal population.
State one reason why the test performed in part (c) might not be valid.
Two additional components are tested at a later date. The mean weight for all twelve components is g and the standard deviation is g.
For further analysis, a 'quality score' for the twelve components is obtained by multiplying the weights by and subtracting .
(e) For the twelve components, find
(i) their mean quality score.
(ii) the standard deviation of their quality score.
Consider how the sample is structured based on characteristics (like production line) and how the specific units are chosen within those structures.
To find the mean, sum all the weights and divide by the number of components.
Use a GDC for efficient calculation of standard deviation. Ensure you are using the sample standard deviation if the context implies the sample is used to estimate a population, or population standard deviation if the sample is the entire population of interest.
Formulate null and alternative hypotheses. Since the population standard deviation is unknown and the sample size is small, a t-test is appropriate. Use your GDC to find the p-value and then compare it to the significance level.
Consider the method used to select the components for testing and whether it truly represents the entire production.
When data is transformed linearly by , the new mean is .
When data is transformed linearly by , the new standard deviation is .
Question 6
HardPaper 2 · calculator18 marksThe battery life, in hours, of a particular smartphone model, , can be modelled by a normal distribution with a mean of 24 hours and a standard deviation of 2 hours.
(a) Find the probability that a randomly selected smartphone has a battery life greater than 27 hours.
Two smartphones are selected at random and independently of each other.
(b) (i) Find the probability that both smartphones have a battery life greater than 27 hours.
(b) (ii) Find the probability that their total battery life is greater than 52 hours.
A software update is released which is claimed to improve battery life. The manufacturer decides to take a random sample of 20 smartphones to test this claim at the 1% significance level, assuming the standard deviation of the battery life has not changed.
(c) Write down the null and alternative hypotheses for the test.
(d) Find the critical region for this test.
Unknown to the manufacturer, the software update has resulted in all smartphones having a 5% longer battery life than the original model.
(e) Find the mean and standard deviation of the battery life for smartphones with the update.
(f) Find the probability of a Type II error in the manufacturer’s test.
Use your GDC's normal distribution function to find the probability for a single smartphone. You are looking for P(L > 27).
The selections are independent. How do you combine probabilities of independent events?
Recall the rules for the mean and variance of the sum of two independent random variables: and . Remember that the standard deviation is the square root of the variance.
The null hypothesis represents 'no change' from the original mean, while the alternative hypothesis represents the manufacturer's claim that the battery life has improved.
The critical region is the set of sample mean values that would lead you to reject the null hypothesis. Find the value `c` such that the probability of the sample mean being greater than `c` is equal to the significance level, under the null hypothesis.
A 5% increase means the new value is 105% of the old value. How does multiplying a random variable by a constant `k` affect its mean and standard deviation?
A Type II error is failing to reject the null hypothesis when it is false. This means the sample mean falls outside the critical region found in part (d). You need to calculate this probability using the true (new) distribution parameters found in part (e).
No question on this page matches those filters. Try another difficulty or paper.
Every Linear transformation of random variable(s), unbiased estimate of mean and var question, marked for you
Every answer is marked mark by mark, IB-style, and the AI tutor helps when you are stuck.
Where marks are lost
- Answering to the wrong accuracy. Two significant figures, or six, where the rule says exactly or three. Common wherever a GDC's full decimal display gets copied straight down.
- Rounding an intermediate value and then using it in a later part. Costs a mark every time, and AI's multi-part modelling questions give it more chances to happen than AA's shorter, more self-contained ones.
- Writing the answer and nothing else, where the mark scheme has an explicit M1 rather than an implied one. A bare answer cannot score full marks there.