Critical Values/Regions, Population Mean Tests (normal/poisson), Proportion Tests (binomial), Correlation Hypothesis Testing, and Type I/II Errors: notes and practice questions
- Hypothesis Test: Uses sample data to evaluate a statement about a population parameter.
- Null Hypothesis (): Default statement (no change, no difference, or no correlation).
- Alternative Hypothesis (): Opposes ().
- p-value: Probability of observing a test statistic at least as extreme as calculated, assuming is true.
- Test Statistic: Value calculated from sample data used for decision.
- Critical Region: Range of test statistic values leading to rejection.
- Critical Value: Boundary of critical region, determined by significance level ().
- Type I Error: Rejecting when it is true.
- Continuous:
- Discrete:
- Type II Error: Failing to reject when it is false.
- Unbiased Estimate of Population Variance:
- Sample Mean Distribution (Z-test, population variance known):
- Conclusion using p-value:
- If p-value < Reject .
- If p-value > Accept .
- Conclusion using Critical Value:
- If test statistic inside critical region Reject .
- If test statistic outside critical region Accept .
- Population Mean Tests:
- Z-Test: Population variance known.
- One-Sample T-Test: Population variance unknown, use .
- Two-Sample T-Test: Compares two population means.
- Binomial Hypothesis Testing (HL): Tests population proportion, always one-tailed. Find critical value using inverse binomial.
- Poisson Hypothesis Testing (HL): Tests mean number of occurrences. Find critical region where cumulative probability .
- Correlation Hypothesis Testing (HL):
- Reject if .
- GDC Use: Utilize built-in hypothesis test functions. For discrete critical regions (HL), manually check cumulative probabilities around inverse function output.
- Z-test standard deviation for sample means:
- HL Topics: Binomial/Poisson testing, critical regions, Type I/II errors, Correlation testing.
- SL Topics: tests, basic t-tests.
- Stating Conclusions: Use "Insufficient evidence to reject " or "Insufficient evidence to suggest [Alternative Hypothesis context]". Do not state is true.
- Poisson Timeframes: Scale mean () to match the test's stated time period.
How it is examined
The hardest statistics content in the course, and a Paper 3 favourite. Critical regions for discrete distributions are the fiddly part: the region is built up one value at a time until adding another would push the Type I error above the significance level, which is why the guide spells out the maximising rule. A Type II error probability needs a specific alternative value to compute against, so the question has to supply one. Matched pairs being a single-sample technique is worth stating on any paired-data question.
- Critical values and critical regions.
- A test for a population mean for a normal distribution.
- A test for a proportion using the binomial distribution.
- A test for a population mean using the Poisson distribution.
**Students will not be expected to calculate critical regions for -tests.**
Linking questions
- Links to other subjects: field studies (sciences and individuals and societies).
- TOK: mathematics and the world. In practical terms, is saying that a result is significant the same as saying that it is true? Does the ability to test only certain parameters in a population affect the way knowledge claims in the human sciences are valued? When is it more important not to make a Type I error and when is it more important not to make a Type II error?
Practice questions
23 questions · 15 medium · 8 hardQuestion 1
MediumPaper 1 · calculator8 marksA renowned artisanal bakery claims that only 5% of its specialty sourdough loaves have minor cosmetic imperfections (e.g., slight cracks, uneven browning). A local restaurant owner, who regularly purchases these loaves, decides to test this claim. For their latest delivery, the owner inspects a batch of 150 loaves and finds 12 loaves with cosmetic imperfections.
(a) Identify the type of sampling used by the restaurant owner.
(b) State the null and alternative hypotheses for this test.
(c) Calculate the p-value for this hypothesis test, assuming cosmetic imperfections occur independently.
(d) The restaurant owner performs the test at the 5% significance level. State the conclusion of the test, giving a reason.
Consider how the sample of loaves was chosen for inspection.
The null hypothesis represents the bakery's claim, while the alternative hypothesis reflects the restaurant owner's suspicion (that the proportion might be higher).
This is a binomial distribution problem. You need to calculate the probability of observing 12 or more imperfect loaves out of 150, given the null hypothesis.
Compare your calculated p-value with the given significance level to determine whether to reject or fail to reject the null hypothesis.
Question 2
HardPaper 2 · calculator15 marksA quality control manager at a manufacturing plant wants to assess the consistency of a new batch of electronic components. He decides to test ten components, ensuring that five are selected from Production Line A and five from Production Line B. The manager instructs the supervisors of each line to provide the required number of components from their current production.
(a) Name the type of sampling that best describes the method used by the quality control manager.
The weights, in grams, of the ten components selected for the test are:
.
(b) For these ten components, find
(i) the mean weight.
(ii) the standard deviation of the weights.
The target weight for these components is g. The manager is concerned that the components might be consistently underweight. Perform an appropriate test at the significance level to see if the mean weight of the components produced is less than the target weight. It can be assumed that the weights come from a normal population.
State one reason why the test performed in part (c) might not be valid.
Two additional components are tested at a later date. The mean weight for all twelve components is g and the standard deviation is g.
For further analysis, a 'quality score' for the twelve components is obtained by multiplying the weights by and subtracting .
(e) For the twelve components, find
(i) their mean quality score.
(ii) the standard deviation of their quality score.
Consider how the sample is structured based on characteristics (like production line) and how the specific units are chosen within those structures.
To find the mean, sum all the weights and divide by the number of components.
Use a GDC for efficient calculation of standard deviation. Ensure you are using the sample standard deviation if the context implies the sample is used to estimate a population, or population standard deviation if the sample is the entire population of interest.
Formulate null and alternative hypotheses. Since the population standard deviation is unknown and the sample size is small, a t-test is appropriate. Use your GDC to find the p-value and then compare it to the significance level.
Consider the method used to select the components for testing and whether it truly represents the entire production.
When data is transformed linearly by , the new mean is .
When data is transformed linearly by , the new standard deviation is .
Question 3
MediumPaper 1 · calculator8 marksA bakery owner, "The Daily Loaf", wants to determine if the sales of their signature "Morning Glory Muffin" are consistent throughout the work week. They believe that the number of muffins sold is the same each day.
To test this, they record the number of muffins sold each weekday during a particular week. This data is shown in the table.
Day | Monday | Tuesday | Wednesday | Thursday | Friday
Number of muffins sold | 120 | 135 | 115 | 140 | 150
A goodness of fit test at the 5% significance level is used on this data to determine whether the owner's belief is suitable.
The critical value for the test is 9.49 and the hypotheses are:
: The number of Morning Glory Muffins sold is consistent across weekdays.
: The number of Morning Glory Muffins sold is not consistent across weekdays.
Find an estimate for how many Morning Glory Muffins the owner expects to sell each day.
Write down the degrees of freedom for this test.
Calculate the statistic for this data.
Using the critical value of 9.49, state the conclusion to the test. Give a reason for your answer.
To find the expected number of sales per day, consider the total sales observed and the number of days.
The degrees of freedom for a goodness of fit test are calculated as the number of categories minus one.
Recall the formula for the chi-squared statistic: , where are the observed frequencies and are the expected frequencies.
Compare your calculated statistic with the given critical value. Remember the significance level and what it means for rejecting or not rejecting the null hypothesis.
Question 4
HardPaper 2 · calculator25 marksThe lifespan of a new type of LED bulb, in hours, can be modelled by a normal distribution with a mean of hours and a standard deviation of hours.
A randomly selected LED bulb is chosen.
(a) Calculate the probability that its lifespan is
(i) less than hours.
(ii) between hours and hours.
(b) of LED bulbs have a lifespan of more than hours.
Calculate the value of .
A manufacturer wants to determine if a sample of LED bulbs from a new production batch could have been chosen from a normally distributed population with a mean of hours and a standard deviation of hours.
They perform a goodness of fit test at the significance level. They begin by creating the following frequency table:
| Lifespan, (hours) | Observed frequency | Expected frequency |
|---|---|---|
| a | ||
| b | ||
(c) Calculate, correct to four significant figures, the value of
(i) a.
(ii) b.
The hypotheses for the manufacturer's test are:
: The lifespans of the LED bulbs are drawn from a normally distributed population with mean hours and standard deviation hours.
: The lifespans of the LED bulbs are not drawn from a normally distributed population with mean hours and standard deviation hours.
(d) Write down the degrees of freedom for this test.
The critical value for this test is .
(e) Perform the goodness of fit test and state your conclusion, justifying your reasoning.
A competitor claims that their new 'Brand A' LED bulbs last longer on average than the manufacturer's 'Brand B' LED bulbs.
Random samples of Brand A bulbs and Brand B bulbs are chosen, and their lifespans (in hours) are measured:
Lifespans of Brand A bulbs (hours):
Lifespans of Brand B bulbs (hours):
The competitor performs a t-test at the significance level. It is assumed that the populations are normally distributed and have equal variances.
(f) Write down the null and alternative hypotheses for this test.
(g) Perform the t-test and state the conclusion, justifying your reasoning.
For a normal distribution, use the normal cumulative distribution function (normal CDF) on your GDC. Remember to input the lower bound, upper bound, mean, and standard deviation.
For the probability between two values, use the normal CDF with the given lower and upper bounds.
If of bulbs last more than hours, then of bulbs last less than or equal to hours. Use the inverse normal function on your GDC.
To find the expected frequency for a category, calculate the probability of a bulb's lifespan falling into that category using the normal distribution, then multiply by the total sample size (). For 'a', calculate .
For 'b', calculate . Alternatively, if you have calculated 'a' and the other expected frequencies, you can subtract them from the total sample size ().
The degrees of freedom for a goodness of fit test are calculated as (number of categories - 1 - number of parameters estimated from the sample). In this case, the mean and standard deviation are given, not estimated.
Use your GDC to perform the GOF test. Compare the p-value to the significance level or the statistic to the critical value.
The null hypothesis () always states that there is no difference or no effect. The alternative hypothesis () reflects the claim being tested, which is that Brand A bulbs last longer on average.
Use your GDC's two-sample t-test function. Ensure you select 'pooled' for equal variances and the correct alternative hypothesis (one-tailed). Compare the p-value to the significance level.
Question 5
MediumPaper 1 · calculator7 marksA bank manager observes that the number of customers arriving at their main ATM during peak hours follows a Poisson distribution with a mean of 18 customers per hour.
To encourage more foot traffic, the bank launches a new promotional campaign. The manager wants to test if this campaign has increased the number of ATM arrivals. They decide to monitor the ATM for a single 4-hour peak period and use a 5% level of significance for their test.
(a) State the null and alternative hypotheses for the test.
(b) Find the probability that the bank manager will make a Type I error in their test conclusion.
During the 4-hour observation period, the bank manager recorded 85 customer arrivals at the ATM.
(c) State the bank manager's conclusion to the test. Justify your answer.
Remember that the null hypothesis represents the status quo or no change, while the alternative hypothesis represents what the manager is trying to prove. Consider the total mean for the observation period.
A Type I error occurs when you incorrectly reject the null hypothesis. This means the observed value falls into the critical region, even though the null hypothesis is true. You need to find the smallest value for which the cumulative probability (or its complement) is less than or equal to the significance level.
Compare the observed number of arrivals to the critical region determined in part (b), or calculate the p-value for the observed number of arrivals and compare it to the significance level.
Question 6
HardPaper 2 · calculator21 marksA company manufactures specialized medical devices. Each device consists of a main circuit board and a protective casing. The weight of the circuit board, , is normally distributed with a mean of g and a standard deviation of g. The weight of the protective casing, , is normally distributed with a mean of g and a standard deviation of g. The weights of the circuit board and the casing are independent.
Find the probability that a randomly chosen complete device has a total weight of less than g.
A batch of such devices is to be packed into a container. The container has a maximum weight capacity of g. The weight of each device is independent.
Find the probability that the total weight of the devices is greater than the capacity of the container.
The company sources a critical microchip component from two different suppliers, Supplier X and Supplier Y. An engineer claims that microchips from Supplier Y have a lower response time than those from Supplier X. To test this claim, a random sample is taken from each supplier.
The eight microchips in the sample from Supplier X have response times, in milliseconds (ms), of:
.
Find the
(i) mean response time for the sample from Supplier X.
Find the
(ii) unbiased estimate of the population variance for the sample from Supplier X.
The seven microchips in the sample from Supplier Y have a mean response time of ms and an unbiased estimate of the population standard deviation () of ms.
Perform a suitable test, at the significance level, to test the engineer's claim that microchips from Supplier Y have a lower response time than those from Supplier X. You may assume the response times of microchips from each supplier are normally distributed with equal population variance.
The total weight of the device is the sum of the circuit board weight and the casing weight. When combining independent normal random variables, their means add, and their variances add. Remember that the standard deviation is the square root of the variance. Once you have the mean and standard deviation of the total weight, you can use the normal distribution to find the required probability.
Let be the total weight of the devices. Since each device's weight is normally distributed and independent, the sum of such weights will also be normally distributed. The mean of the sum will be times the mean of a single device, and the variance of the sum will be times the variance of a single device. Use the mean and variance calculated in part (a).
To find the mean of a sample, sum all the values in the sample and divide by the number of values in the sample.
The unbiased estimate of the population variance, , is calculated using the formula , where is the sample size and is the sample mean. Be careful to use in the denominator.
This is a hypothesis test comparing the means of two independent samples. Since the population standard deviations are unknown but assumed equal, and the samples are from normally distributed populations, a two-sample t-test (pooled variance) is appropriate. Remember to state the null and alternative hypotheses, calculate the test statistic and p-value, and draw a conclusion in context based on the significance level.
Question 7
MediumPaper 1 · calculator6 marksA materials scientist is developing a new synthetic fiber and claims its average tensile strength is 120 N. To test this claim, she takes a random sample of 25 fiber segments.
Given that the tensile strength of an individual fiber segment is newtons, the scientist found that, from her 25 samples, and .
(a) Find an unbiased estimate for the mean tensile strength () of the fiber.
(b) Use the formula to determine an unbiased estimate for the variance of the tensile strength of the fiber.
(c) Find a 95% confidence interval for . You may assume that all conditions for a confidence interval have been met.
(d) Suggest, with justification, a valid conclusion that the materials scientist could make regarding her claim.
Recall the formula for the sample mean, which is an unbiased estimate for the population mean.
Carefully substitute the given values into the formula for the unbiased sample variance. Remember to use in the denominator.
You will need the sample mean and the unbiased standard deviation (square root of the variance) from parts (a) and (b). For a 95% confidence interval with a small sample size and unknown population standard deviation, use the t-distribution. You can use your GDC to find the critical t-value or the entire confidence interval.
Compare the claimed average tensile strength (120 N) with the 95% confidence interval you calculated in part (c). If the claimed value falls outside the interval, what does that imply?
Question 8
HardPaper 2 · calculator18 marks(a) The manager of "The Daily Grind" coffee shop suggested that the number of customers arriving at the shop during a 5-minute interval can be modelled by a Poisson distribution.
Suggest two observations that the manager may have made that led him to suggest this model.
(b) Now assume that the model is valid and that the mean number of customers arriving at the shop during a 5-minute interval is .
The manager observes customer arrivals during a 15-minute interval.
Calculate the probability that exactly 6 customers arrive during this 15-minute interval.
(c) Using the same model as in part (b), find the probability that fewer than 4 customers arrive during a 15-minute interval.
(d) Find the probability that in four consecutive 5-minute intervals, at least one customer arrives in each interval.
(e) Following a new marketing campaign, the manager wished to determine whether the mean number of customers arriving during a 5-minute interval had increased.
State the hypotheses for the test.
(f) Find the critical region for the test at the 5% significance level.
(g) Given that the mean number of customers per 5-minute interval has actually risen to , find the probability that the manager makes a Type II error.
Recall the key assumptions of a Poisson distribution regarding the nature of events.
First, adjust the mean () for the new time interval. Then, use the Poisson probability mass function .
Remember that 'fewer than 4' means , which is equivalent to . Use the Poisson cumulative distribution function.
First, calculate the probability of at least one customer arriving in a single 5-minute interval. Then, consider how to combine probabilities for independent events.
Formulate the null hypothesis () as no change, and the alternative hypothesis () reflecting the manager's suspicion of an increase. Use the notation for the mean of a Poisson distribution.
For a one-tailed test for an increase, you need to find a value such that the probability of observing or more customers (under the null hypothesis) is less than or equal to the significance level.
A Type II error occurs when you fail to reject the null hypothesis when it is false. This means observing a value outside the critical region, given the true mean is .
Question 9
MediumPaper 1 · calculator7 marksThe number of incoming calls received by a tech support hotline in a particular 30-minute period is historically known to follow a Poisson distribution with a mean of 48 calls per hour.
A new automated system is implemented, and it is claimed that this system has decreased the number of calls requiring human intervention.
To test this claim, the number of calls, Y, requiring human intervention in a 30-minute period on a particular day will be recorded. The test will have the following hypotheses:
: the mean number of calls requiring human intervention has not changed,
: the mean number of calls requiring human intervention has decreased.
The alternative hypothesis will be accepted if .
Assuming the null hypothesis to be true, state the distribution of Y.
Find the probability of a Type I error.
Find the probability of a Type II error, if the number of calls now follows a Poisson distribution with a mean of 35 calls per hour.
Recall the definition of a Poisson distribution and how to calculate its mean over a specific time period.
A Type I error occurs when the null hypothesis is rejected when it is actually true. This corresponds to the probability of observing a value in the critical region under the null hypothesis.
A Type II error occurs when the null hypothesis is accepted when it is actually false. This means observing a value outside the critical region when the alternative hypothesis is true. Remember to calculate the new mean for the observation period.
Question 10
HardPaper 3 · calculator24 marksMs. Anya Sharma, a school principal, wants to investigate if the number of hours students spend studying affects their exam scores. This question asks you to review Ms. Sharma's methods and conclusions.
Ms. Sharma obtained a list of students from her school. She contacted them and asked them to fill in an anonymous questionnaire. Participants were asked to state their weekly study hours and their most recent exam score (out of 100). Of the 250 students on the list, 11 replied.
Ms. Sharma's results are summarized in the following table:
| Student ID | Weekly Study Hours (X) | Exam Score (Y) |
|---|---|---|
| 1 | 5 | 50 |
| 2 | 7 | 65 |
| 3 | 8 | 60 |
| 4 | 10 | 78 |
| 5 | 12 | 70 |
| 6 | 6 | 55 |
| 7 | 9 | 72 |
| 8 | 11 | 80 |
| 9 | 4 | 45 |
| 10 | 13 | 85 |
| 11 | 18 | 60 |
Describe one way in which Ms. Sharma could improve the reliability of her investigation.
Describe one criticism that can be made about the validity of Ms. Sharma's investigation.
Ms. Sharma classifies Student 11 as an outlier and removes their data from the analysis. Suggest one possible justification for her decision to remove it.
For the remaining ten student responses in the table, Ms. Sharma calculates the mean exam score to be . Calculate the mean weekly study hours for these remaining responses.
Determine the value of , Pearson's product-moment correlation coefficient, for these remaining responses.
Ms. Sharma decides to carry out a hypothesis test on the correlation coefficient to investigate whether increased weekly study hours are associated with higher exam scores. State why the hypothesis test should be one-tailed.
State the null and alternative hypotheses for this test.
The critical value for this test, at the 5% significance level, is 0.549. Ms. Sharma assumes that the population is bivariate normal. Determine whether there is significant evidence of a positive correlation between weekly study hours and exam scores. Justify your answer.
Ms. Sharma wants to create a model to predict how changing weekly study hours might affect exam scores. To do this, she assumes that weekly study hours, , is the independent variable and the exam score, , is the dependent variable.
She first considers a linear model of the form . Use Ms. Sharma's data to find the value of and of .
Interpret, referring to study hours and exam scores, what the value of represents.
Ms. Sharma then considers a quadratic model of the form . Find the value of , of and of .
Find the coefficient of determination for each of the two models she considers.
Hence compare the two models.
Ms. Sharma decides to use the coefficient of determination to choose between these two models. Comment on the validity of her decision.
After presenting the results of her investigation, a colleague questions whether Ms. Sharma's sample is representative of all students in the school. A report states that the mean weekly study hours for all students in the school is hours. Ms. Sharma decides to carry out a test to determine whether her sample could realistically be taken from a population with a mean of hours. State the name of the test which Ms. Sharma should use.
State the null and alternative hypotheses for this test.
Perform the test, using a 5% significance level, and state your conclusion in context.
Reliability concerns the consistency and repeatability of the results. How can she ensure her measurements are more consistent or less prone to random error?
Validity concerns whether the study measures what it intends to measure and whether the results are generalizable. Are there other factors influencing exam scores? Is "study hours" accurately measured?
Look at the data for Student 11 compared to the general trend. What makes it unusual?
Sum the weekly study hours for the remaining students and divide by .
Use your GDC's statistical functions to calculate Pearson's for the data points (excluding Student 11).
Consider the specific direction of the relationship Ms. Sharma is investigating.
The null hypothesis typically states no effect or no relationship, while the alternative hypothesis states the effect or relationship you are looking for. Use the correct symbol for population correlation.
Compare your calculated value from part (c.ii) with the given critical value.
Use your GDC's linear regression function (LinReg(ax+b) ) with the data points.
The coefficient in a linear model represents the change in for every one-unit increase in .
Use your GDC's quadratic regression function (QuadReg) with the data points.
The coefficient of determination, , is often provided by your GDC along with the regression equation. For the linear model, .
A higher value generally indicates a better fit for the data.
tends to increase with the number of independent variables or parameters in a model, even if the additional terms do not significantly improve the model's predictive power.
This is a test comparing a sample mean to a known population mean when the population standard deviation is unknown (which is usually the case).
The null hypothesis assumes the sample comes from the population with the stated mean. The alternative hypothesis states it does not.
Use your GDC's t-test function (T-Test) for one sample. Input the sample data (study hours), the hypothesized population mean, and the significance level.
Question 11
MediumPaper 1 · calculator6 marksA factory produces two types of electronic components: standard (S) and premium (P). The weight of these components is a critical characteristic for quality control.
The weights of standard components are known to be normally distributed with a mean of 150 grams and a standard deviation of 5 grams.
The weights of premium components are known to be normally distributed with a mean of 165 grams and a standard deviation of 8 grams.
A quality control machine classifies a component as 'premium' if its weight is found to be above 158 grams; otherwise, it is classified as 'standard'.
The factory's quality control manager uses the null hypothesis that, in the absence of other information, a component is standard.
Calculate the probability of making a Type I error when classifying a component.
Calculate the probability of making a Type II error when classifying a component.
It is known that 80% of the components produced are standard, and 20% are premium.
Calculate the overall probability that a randomly selected component is misclassified by the machine.
A Type I error occurs when the null hypothesis is true, but it is rejected. In this context, consider which type of component is incorrectly classified as the other.
A Type II error occurs when the null hypothesis is false, but it is not rejected. In this context, consider which type of component is incorrectly classified as the other.
Consider the total probability of error by combining the probabilities of Type I and Type II errors with the prior probabilities of each component type.
Question 12
HardPaper 3 · calculator27 marks(a) Mr. Lee, the owner of "Sweet Delights" bakery, recorded the number of Mooncakes sold each day for a sample of days. The results are shown in the table below.
| Number of Mooncakes sold | Frequency |
|---|---|
| 0 | 1 |
| 1 | 2 |
| 2 | 4 |
| 3 | 5 |
| 4 | 7 |
| 5 | 6 |
| 6 | 3 |
| 7 | 2 |
(a.i) Find the mean and variance for this sample data.
(a.ii) Hence, state why Mr. Lee might believe that the daily sales of Mooncakes follow a Poisson distribution.
(b) State one assumption that Mr. Lee needs to make about the sales of Mooncakes to support his belief that it follows a Poisson distribution.
(c) Mr. Lee knows from his historic sales records that the bakery sells an average of Mooncakes each day. The following table shows the expected frequency of Mooncakes sold each day during a -day period, assuming a Poisson distribution with mean .
| Number of Mooncakes sold | <1 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | |
|---|---|---|---|---|---|---|---|---|---|
| Expected frequency | a | b | c |
Find the value of a, of b, and of c. Give your answers to 3 decimal places.
(d) Mr. Lee decides to carry out a goodness of fit test at the significance level to see whether the daily sales of Mooncakes follow a Poisson distribution with mean . He collects observed frequencies for days, which are given in the table below.
| Number of Mooncakes sold | <2 | 2 | 3 | 4 | 5 | 6 | |
|---|---|---|---|---|---|---|---|
| Observed frequency | 12 | 13 | 23 | 17 | 12 | 10 | 13 |
| Expected frequency |
(d.i) Write down the number of degrees of freedom for his test.
(d.ii) Perform the goodness of fit test and state, with reason, a conclusion.
(e) Mr. Lee claims that a new social media advertising campaign, costing THB per day, will increase the number of Mooncakes sold. However, his business partner, Ms. Chen, claims that the advertising will not increase the bakery's overall profit.
Ms. Chen agrees to run the campaign for the next days. During that time, Mr. Lee records that the bakery sells a total of Mooncakes, with a profit of THB on each Mooncake sold.
Mr. Lee wants to carry out an appropriate hypothesis test to determine whether the number of Mooncakes sold during the days increased when compared with the historic sales records (mean Mooncakes per day).
By finding a critical value, perform this test at a significance level.
(f) Hence state the probability of a Type I error for this test.
(g) By considering the claims of both Mr. Lee and Ms. Chen, explain whether the advertising campaign was beneficial to the bakery.
To find the mean, calculate the sum of (number of mooncakes frequency) and divide by the total frequency. For the variance, use the formula for sample variance, .
Recall the relationship between the mean and variance for a Poisson distribution.
Consider the characteristics of events that follow a Poisson distribution, such as independence or constant rate.
Use the Poisson probability mass function to find the probabilities for each category. Multiply these probabilities by the total number of days () to get the expected frequencies. For '', calculate .
The degrees of freedom for a goodness of fit test are , where is the number of categories and is the number of parameters estimated from the data. In this case, the mean is given.
First, state the null and alternative hypotheses. Then, calculate the test statistic using the formula . Use your GDC to find the p-value for the calculated test statistic and degrees of freedom. Compare the p-value to the significance level to draw a conclusion.
First, calculate the expected total number of Mooncakes sold over 40 days based on the historic mean. Define your null and alternative hypotheses. Since this is a Poisson distribution, you need to find the critical value such that , where follows a Poisson distribution with the expected total mean. Compare the observed total sales with this critical value.
The probability of a Type I error is the significance level of the test, specifically, the probability of rejecting the null hypothesis when it is actually true. This is the probability of observing a result as extreme as, or more extreme than, the critical value, assuming the null hypothesis is true.
Calculate the total cost of the advertising campaign and the additional profit generated from the increased sales. Compare these two values to determine the overall financial impact.
Question 13
MediumPaper 1 · calculator8 marksA manufacturing company produces electronic components. Historically, the defect rate for a specific component has been 15%. A new production line is implemented, and the quality control manager wants to test if the defect rate has decreased. She assumes that the defect status of each component is independent of others.
(a) Write down suitable hypotheses for this test.
(b) The quality control manager decides to take a random sample of 120 components. She will reject the null hypothesis if fewer than 12 components are found to be defective.
Find the probability that she makes a Type I error.
(c) In fact, the new production line successfully reduced the defect rate to 10%.
Find the probability that she makes a Type II error.
Remember to define your parameter (e.g., p for probability) and state both the null and alternative hypotheses using appropriate notation. Consider what 'decreased' implies for the alternative hypothesis.
A Type I error occurs when you reject the null hypothesis () when it is actually true. Use the binomial distribution with the parameters defined by and the rejection region given.
A Type II error occurs when you accept the null hypothesis () when it is false. This means using the true defect rate (10%) and the acceptance region (the complement of the rejection region from part (b) ).
Question 14
HardPaper 3 · calculator28 marks(a) TechInnovate is considering collecting more data for their analysis.
(i) State one advantage of increasing the sample size.
(ii) State one disadvantage of increasing the sample size.
(b) The production manager at Plant Alpha recorded the time, in minutes, taken to produce a batch of electronic components for 10 randomly selected batches:
Find the value of for this sample from Plant Alpha.
(c) A manager claims that Plant Alpha's production times are more consistent than Plant Beta's. Given that the sample standard deviation () for Plant Beta's production times is minutes, make one criticism of this claim.
(d) TechInnovate wants to compare the mean production times of Plant Alpha and Plant Beta using a pooled t-test.
(i) State the condition regarding population variances required to use a pooled t-test.
(ii) Given that for Plant Beta, a sample of batches yielded a mean production time of minutes and a sample standard deviation of minutes, state whether TechInnovate should use a pooled t-test in this case. Justify your answer.
(e) TechInnovate believes Plant Alpha has a lower mean production time than Plant Beta.
(i) State appropriate null and alternative hypotheses for the pooled t-test.
(ii) Find the p-value.
(iii) Given that the test is carried out at the 5% significance level, state the appropriate conclusion in context. Justify your answer.
(f) The company also investigates the relationship between operator experience (in years) and the number of defective items produced per day. A sample of 8 operators yielded the following data:
| Operator Experience (years) | Number of Defective Items |
|---|---|
| 2 | 15 |
| 5 | 10 |
| 3 | 13 |
| 8 | 7 |
| 1 | 18 |
| 6 | 9 |
| 4 | 12 |
| 7 | 8 |
(i) Assuming all requirements are met, perform a test at the 5% significance level to determine if there is a linear correlation between operator experience and the number of defective items. State the hypotheses and justify your conclusion.
(ii) If the requirements for this test are not met, state an alternative test that could be used.
(g) For the data in (f.i), the equation of the least squares regression line of defective items () on operator experience () is . Give, in context, an interpretation of the gradient in this model.
(h) TechInnovate uses a baseline model to predict the number of defective items () for a batch based on its size (): . The "Quality Deviation" () for a batch is defined as . A positive Quality Deviation indicates better-than-expected quality.
(i) Show that for a batch of components from Plant Beta that produced defective items, the Quality Deviation is .
(ii) To compare quality control, samples of Quality Deviation scores were collected:
- Plant Alpha: , ,
- Plant Beta: , ,
Assuming that the appropriate requirements are met, use a pooled t-test at a 5% significance level to determine if the mean Quality Deviation is higher in Plant Alpha than in Plant Beta. Write down your null and alternative hypotheses and justify your conclusion.
(i) Using the results from parts (e) and (h.ii), explain how each plant could claim they are performing better than the other plant.
Think about how a larger sample relates to the overall population.
Consider the practical implications of collecting more data.
Use your GDC to calculate the sample standard deviation ( or ).
Consider the nature of sample statistics versus population parameters, especially when values are close.
Recall the assumption about variances for a pooled t-test.
Compare the sample standard deviations of Plant Alpha (from part b) and Plant Beta.
Remember to define your parameters and specify the direction of the alternative hypothesis.
Use your GDC to perform a two-sample t-test with pooled variance, then adjust for the one-tailed hypothesis.
Compare the p-value with the significance level and relate it back to the original claim about production times.
Calculate the Pearson product-moment correlation coefficient () and its associated p-value. Formulate hypotheses for population correlation ().
Consider non-parametric alternatives for correlation when assumptions for Pearson's are violated.
The gradient represents the change in the dependent variable for a one-unit change in the independent variable.
First, calculate the predicted number of defective items using the model. Then, apply the definition of Quality Deviation.
Formulate hypotheses for the population mean Quality Deviation. Perform a one-tailed pooled t-test using your GDC.
Review the conclusions of the two t-tests. One test might favor Plant Alpha, while the other might not show a significant difference, which Plant Beta could use to their advantage.
Question 15
MediumPaper 1 · calculator6 marksMs. Chen, a small business owner, wants to investigate if there is a monotonic relationship between the monthly advertising budget and the number of units of a new product sold. She collects data for eight months, as shown in Table 1.
Table 1: Monthly Advertising Budget and Units Sold
| Month | Advertising Budget (in $100s) | Units Sold |
|---|---|---|
| 1 | 2.5 | 120 |
| 2 | 3.0 | 150 |
| 3 | 1.8 | 95 |
| 4 | 4.2 | 170 |
| 5 | 3.5 | 110 |
| 6 | 2.0 | 180 |
| 7 | 4.8 | 230 |
| 8 | 3.2 | 130 |
Ms. Chen decides to calculate the Spearman's rank correlation coefficient. Complete the table of ranks shown in Table 2.
Table 2: Ranks for Advertising Budget and Units Sold
| Month | Rank of Advertising Budget () | Rank of Units Sold () |
|---|---|---|
| 1 | 3 | 3 |
| 2 | 5 | |
| 3 | 1 | 1 |
| 4 | 7 | |
| 5 | 2 | |
| 6 | 2 | |
| 7 | 8 | 8 |
| 8 | 4 |
Calculate the value of , Spearman's rank correlation coefficient.
Ms. Chen believes that a higher advertising budget leads to more units sold. She carries out a hypothesis test using a 10% significance level with the following null hypothesis:
: In the population, there is no monotonic relationship between the monthly advertising budget and the number of units sold.
Write down Ms. Chen's alternative hypothesis.
The critical value of for this test is 0.643.
State the conclusion of the hypothesis test, giving a reason.
To rank the data, assign rank 1 to the smallest value, rank 2 to the next smallest, and so on. If there are tied values, assign the average of the ranks they would have occupied.
Recall the formula for Spearman's rank correlation coefficient: , where is the difference in ranks for each pair of data and is the number of data pairs.
The alternative hypothesis () should reflect Ms. Chen's belief and contradict the null hypothesis. Consider if it's a positive, negative, or non-directional relationship.
Compare your calculated value from part (b) with the given critical value . If , you reject . Otherwise, you do not reject . Remember to consider the direction of the test.
Question 16
HardPaper 2 · calculator18 marksThe battery life, in hours, of a particular smartphone model, , can be modelled by a normal distribution with a mean of 24 hours and a standard deviation of 2 hours.
(a) Find the probability that a randomly selected smartphone has a battery life greater than 27 hours.
Two smartphones are selected at random and independently of each other.
(b) (i) Find the probability that both smartphones have a battery life greater than 27 hours.
(b) (ii) Find the probability that their total battery life is greater than 52 hours.
A software update is released which is claimed to improve battery life. The manufacturer decides to take a random sample of 20 smartphones to test this claim at the 1% significance level, assuming the standard deviation of the battery life has not changed.
(c) Write down the null and alternative hypotheses for the test.
(d) Find the critical region for this test.
Unknown to the manufacturer, the software update has resulted in all smartphones having a 5% longer battery life than the original model.
(e) Find the mean and standard deviation of the battery life for smartphones with the update.
(f) Find the probability of a Type II error in the manufacturer’s test.
Use your GDC's normal distribution function to find the probability for a single smartphone. You are looking for P(L > 27).
The selections are independent. How do you combine probabilities of independent events?
Recall the rules for the mean and variance of the sum of two independent random variables: and . Remember that the standard deviation is the square root of the variance.
The null hypothesis represents 'no change' from the original mean, while the alternative hypothesis represents the manufacturer's claim that the battery life has improved.
The critical region is the set of sample mean values that would lead you to reject the null hypothesis. Find the value `c` such that the probability of the sample mean being greater than `c` is equal to the significance level, under the null hypothesis.
A 5% increase means the new value is 105% of the old value. How does multiplying a random variable by a constant `k` affect its mean and standard deviation?
A Type II error is failing to reject the null hypothesis when it is false. This means the sample mean falls outside the critical region found in part (d). You need to calculate this probability using the true (new) distribution parameters found in part (e).
Question 17
MediumPaper 1 · calculator6 marksA botanist is investigating the effectiveness of two different soil compositions, Soil A and Soil B, on the growth of a particular plant species. They hypothesize that Soil A will lead to a greater mean plant height after four weeks compared to Soil B. The botanist grows a random sample of plants in each soil type and measures their heights (in cm) after four weeks. The results are shown in the table below.
| Soil Type | Plant Heights (cm) |
|---|---|
| Soil A | 25.5, 23.9, 26.6, 27.2, 25.1, 24.5, 25.4, 25.9, 24.8, 26.1 |
| Soil B | 22.3, 23.4, 22.8, 23.2, 23.7, 24.1, 22.9, 23.5, 23.1, 22.5, 23.8, 23.0 |
The botanist performs a one-tailed t-test at a 5% level of significance. It is assumed that the plant heights are normally distributed and the samples have equal variances.
State the null and alternative hypotheses.
Calculate the p-value for this test.
State the conclusion of the test in the context of the question. Justify your answer.
Remember to define your population means clearly for each soil type. The alternative hypothesis should reflect the botanist's belief about which soil is better.
Use your GDC's statistical test function for a two-sample t-test. Remember to select the correct alternative hypothesis for a one-tailed test and ensure 'pooled' is set to 'yes' for equal variances.
Compare your calculated p-value to the given significance level (5%). Based on this comparison, decide whether to reject or fail to reject the null hypothesis, and then interpret this decision in terms of the plant heights and soil types.
Question 18
MediumPaper 1 · calculator7 marksA company, "Electro-Tech", manufactures electronic components. Their quality control department claims that the average number of defective components in a standard batch of 1000 units follows a Poisson distribution with a mean of defects per batch. A new supplier's components are being tested. To assess their quality, standard batches of components from the new supplier are randomly selected and inspected. A total of defective components are found across these batches.
Test the claim that the new supplier's components have an average of defects per batch against the suspicion that they have more defects, at the significance level. In your answer, you should include whether this is a one-tailed or two-tailed test, the test hypotheses, calculation of the -value, and the conclusion (with a reason) of the test.
Remember that if the mean rate for a single batch is , then the mean rate for batches is . For a Poisson distribution, the p-value for a 'greater than' alternative hypothesis is calculated as . Use the cumulative distribution function (CDF) for the Poisson distribution to find this probability.
Question 19
MediumPaper 2 · calculator13 marksThe battery life, in hours, of eight Brand A batteries was recorded as .
The battery life, in hours, of Brand B batteries was recorded as .
(a) Find the sample mean battery life for
(i) Brand A batteries.
(ii) Brand B batteries.
You can assume that both sets of values have a common unknown variance.
(b) Carry out a test at the significance level to determine if the population mean battery life of Brand A is longer than that of Brand B. In your answer, you should include the test used (with a reason), the test hypotheses, the -value and the conclusion of the test (with a reason).
(c) State if your conclusion would have been any different if working at the significance level.
To find the sample mean, sum all the values in the sample and divide by the number of values in that sample.
To find the sample mean, sum all the values in the sample and divide by the number of values in that sample.
This involves a hypothesis test for the difference between two population means. Consider the conditions for using a t-test and how to formulate one-tailed hypotheses.
Compare the calculated -value from part (b) with the new significance level.
Question 20
MediumPaper 1 · calculator5 marksA global tech company has historically received customer complaints for its flagship software product following a Poisson distribution with a mean of per week. Recently, after a major software update, the company's quality assurance team wants to investigate if the number of complaints has increased.
Over a period of weeks following the update, they recorded a total of complaints.
Test at the significance level the hypothesis that the mean number of complaints has increased.
Start by defining your null and alternative hypotheses. Remember to adjust the mean of the Poisson distribution for the observation period (3 weeks) under the null hypothesis. Then, calculate the probability of observing at least 35 complaints given this adjusted mean, and compare it to the significance level.
No question on this page matches those filters. Try another difficulty or paper.
3 more Critical Values/Regions, Population Mean Tests (normal/poisson), Proportion Tests (binomial), Correlation Hypothesis Testing, and Type I/II Errors questions in the app
Every answer is marked mark by mark, IB-style, and the AI tutor helps when you are stuck.
Where marks are lost
- Rounding an intermediate value and then using it in a later part. Costs a mark every time, and AI's multi-part modelling questions give it more chances to happen than AA's shorter, more self-contained ones.
- Using your own wrong value after failing a "show that." All follow through is withdrawn for the rest of that question.
- Leaving an answer in calculator notation. Never accepted in a final answer, and AI's constant calculator use makes this the easiest slip in the whole subject.