Skip to content
  1. IB Question Bank
  2. Maths AI
  3. Statistics & Probability
Topic 4.19 · HL only

Critical Values/Regions, Population Mean Tests (normal/poisson), Proportion Tests (binomial), Correlation Hypothesis Testing, and Type I/II Errors: notes and practice questions

Summary
  • Hypothesis Test: Uses sample data to evaluate a statement about a population parameter.
  • Null Hypothesis (H0H_0): Default statement (no change, no difference, or no correlation).
  • Alternative Hypothesis (H1H_1): Opposes H0H_0 (>,<, or ≠>, <, \text{ or } \neq).
  • p-value: Probability of observing a test statistic at least as extreme as calculated, assuming H0H_0 is true.
  • Test Statistic: Value calculated from sample data used for decision.
  • Critical Region: Range of test statistic values leading to H0H_0 rejection.
  • Critical Value: Boundary of critical region, determined by significance level (α\alpha).
  • Type I Error: Rejecting H0H_0 when it is true.
  • Continuous: P(Type I error)=αP(\text{Type I error}) = \alpha
  • Discrete: P(Type I error)≤αP(\text{Type I error}) \leq \alpha
  • Type II Error: Failing to reject H0H_0 when it is false.
  • P(Type II error)=P(not in critical region∣actual population parameter)P(\text{Type II error}) = P(\text{not in critical region} | \text{actual population parameter})
  • Unbiased Estimate of Population Variance:

sn−12=nn−1sn2 s_{n-1}^2 = \frac{n}{n-1} s_n^2

  • Sample Mean Distribution (Z-test, population variance σ2\sigma^2 known):

X‾∼N(μ,σ2n) \overline{X} \sim N\left(\mu, \frac{\sigma^2}{n}\right)

  • Conclusion using p-value:
  • If p-value < α⇒\alpha \Rightarrow Reject H0H_0.
  • If p-value > α⇒\alpha \Rightarrow Accept H0H_0.
  • Conclusion using Critical Value:
  • If test statistic inside critical region ⇒\Rightarrow Reject H0H_0.
  • If test statistic outside critical region ⇒\Rightarrow Accept H0H_0.
  • Population Mean Tests:
  • Z-Test: Population variance known.
  • One-Sample T-Test: Population variance unknown, use sn−12s_{n-1}^2.
  • Two-Sample T-Test: Compares two population means.
  • Binomial Hypothesis Testing (HL): Tests population proportion, always one-tailed. Find critical value using inverse binomial.
  • Poisson Hypothesis Testing (HL): Tests mean number of occurrences. Find critical region where cumulative probability ≤α\leq \alpha.
  • Correlation Hypothesis Testing (HL):
  • H0:ρ=0H_0: \rho = 0
  • H1:ρ>0,ρ<0, or ρ≠0H_1: \rho > 0, \rho < 0, \text{ or } \rho \neq 0
  • Reject H0H_0 if ∣PMCC∣>∣criticalvalue∣|PMCC| > |critical value|.
  • GDC Use: Utilize built-in hypothesis test functions. For discrete critical regions (HL), manually check cumulative probabilities around inverse function output.
  • Z-test standard deviation for sample means: σn\frac{\sigma}{\sqrt{n}}
  • HL Topics: Binomial/Poisson testing, critical regions, Type I/II errors, Correlation ρ\rho testing.
  • SL Topics: χ2\chi^2 tests, basic t-tests.
  • Stating Conclusions: Use "Insufficient evidence to reject H0H_0" or "Insufficient evidence to suggest [Alternative Hypothesis context]". Do not state H0H_0 is true.
  • Poisson Timeframes: Scale mean (mm) to match the test's stated time period.

How it is examined

The hardest statistics content in the course, and a Paper 3 favourite. Critical regions for discrete distributions are the fiddly part: the region is built up one value at a time until adding another would push the Type I error above the significance level, which is why the guide spells out the maximising rule. A Type II error probability needs a specific alternative value to compute against, so the question has to supply one. Matched pairs being a single-sample technique is worth stating on any paired-data question.

Key ideas
  • Critical values and critical regions.
  • A test for a population mean for a normal distribution.
  • A test for a proportion using the binomial distribution.
  • A test for a population mean using the Poisson distribution.
Not assessed

**Students will not be expected to calculate critical regions for tt-tests.**

Linking questions

  • Links to other subjects: field studies (sciences and individuals and societies).
  • TOK: mathematics and the world. In practical terms, is saying that a result is significant the same as saying that it is true? Does the ability to test only certain parameters in a population affect the way knowledge claims in the human sciences are valued? When is it more important not to make a Type I error and when is it more important not to make a Type II error?

Practice questions

23 questions · 15 medium · 8 hard
Showing 20 of 20

Question 1

MediumPaper 1 · calculator8 marks
(a)

A renowned artisanal bakery claims that only 5% of its specialty sourdough loaves have minor cosmetic imperfections (e.g., slight cracks, uneven browning). A local restaurant owner, who regularly purchases these loaves, decides to test this claim. For their latest delivery, the owner inspects a batch of 150 loaves and finds 12 loaves with cosmetic imperfections.

(a) Identify the type of sampling used by the restaurant owner.

[1]
(b)

(b) State the null and alternative hypotheses for this test.

[2]
(c)

(c) Calculate the p-value for this hypothesis test, assuming cosmetic imperfections occur independently.

[3]
(d)

(d) The restaurant owner performs the test at the 5% significance level. State the conclusion of the test, giving a reason.

[2]

Question 2

HardPaper 2 · calculator15 marks
(a)

A quality control manager at a manufacturing plant wants to assess the consistency of a new batch of electronic components. He decides to test ten components, ensuring that five are selected from Production Line A and five from Production Line B. The manager instructs the supervisors of each line to provide the required number of components from their current production.

(a) Name the type of sampling that best describes the method used by the quality control manager.

[1]
(b)(i)

The weights, in grams, of the ten components selected for the test are:

148,153,161,155,142,160,150,157,145,159148, 153, 161, 155, 142, 160, 150, 157, 145, 159.

(b) For these ten components, find

(i) the mean weight.

[2]
(b)(ii)

(ii) the standard deviation of the weights.

[2]
(c)

The target weight for these components is 155155 g. The manager is concerned that the components might be consistently underweight. Perform an appropriate test at the 10%10\% significance level to see if the mean weight of the components produced is less than the target weight. It can be assumed that the weights come from a normal population.

[5]
(d)

State one reason why the test performed in part (c) might not be valid.

[1]
(e)(i)

Two additional components are tested at a later date. The mean weight for all twelve components is 154.5154.5 g and the standard deviation is 7.27.2 g.

For further analysis, a 'quality score' for the twelve components is obtained by multiplying the weights by 1.51.5 and subtracting 5050.

(e) For the twelve components, find

(i) their mean quality score.

[2]
(e)(ii)

(ii) the standard deviation of their quality score.

[2]

Question 3

MediumPaper 1 · calculator8 marks
(a)

A bakery owner, "The Daily Loaf", wants to determine if the sales of their signature "Morning Glory Muffin" are consistent throughout the work week. They believe that the number of muffins sold is the same each day.

To test this, they record the number of muffins sold each weekday during a particular week. This data is shown in the table.

Day | Monday | Tuesday | Wednesday | Thursday | Friday

Number of muffins sold | 120 | 135 | 115 | 140 | 150

A goodness of fit test at the 5% significance level is used on this data to determine whether the owner's belief is suitable.

The critical value for the test is 9.49 and the hypotheses are:

H0H_0: The number of Morning Glory Muffins sold is consistent across weekdays.

H1H_1: The number of Morning Glory Muffins sold is not consistent across weekdays.

Find an estimate for how many Morning Glory Muffins the owner expects to sell each day.

[1]
(b)(i)

Write down the degrees of freedom for this test.

[1]
(b)(ii)

Calculate the χ2\chi^2 statistic for this data.

[3]
(b)(iii)

Using the critical value of 9.49, state the conclusion to the test. Give a reason for your answer.

[3]

Question 4

HardPaper 2 · calculator25 marks
(a)(i)

The lifespan of a new type of LED bulb, in hours, can be modelled by a normal distribution with a mean of 1200012000 hours and a standard deviation of 800800 hours.

A randomly selected LED bulb is chosen.

(a) Calculate the probability that its lifespan is

(i) less than 1100011000 hours.

[4]
(a)(ii)

(ii) between 1150011500 hours and 1250012500 hours.

[4]
(b)

(b) 15%15\% of LED bulbs have a lifespan of more than hh hours.

Calculate the value of hh.

[2]
(c)(i)

A manufacturer wants to determine if a sample of 250250 LED bulbs from a new production batch could have been chosen from a normally distributed population with a mean of 1200012000 hours and a standard deviation of 800800 hours.

They perform a χ2\chi^2 goodness of fit test at the 5%5\% significance level. They begin by creating the following frequency table:

Lifespan, hh (hours)Observed frequencyExpected frequency
h≤11000h \le 11000151526.41226.412
11000<h≤1200011000 < h \le 12000100100a
12000<h≤1300012000 < h \le 13000110110b
h>13000h > 13000252526.41226.412

(c) Calculate, correct to four significant figures, the value of

(i) a.

[2]
(c)(ii)

(ii) b.

[2]
(d)

The hypotheses for the manufacturer's test are:

H0H_0: The lifespans of the LED bulbs are drawn from a normally distributed population with mean 1200012000 hours and standard deviation 800800 hours.

H1H_1: The lifespans of the LED bulbs are not drawn from a normally distributed population with mean 1200012000 hours and standard deviation 800800 hours.

(d) Write down the degrees of freedom for this test.

The critical value for this test is 7.8157.815.

[1]
(e)

(e) Perform the χ2\chi^2 goodness of fit test and state your conclusion, justifying your reasoning.

[4]
(f)

A competitor claims that their new 'Brand A' LED bulbs last longer on average than the manufacturer's 'Brand B' LED bulbs.

Random samples of 1212 Brand A bulbs and 1010 Brand B bulbs are chosen, and their lifespans (in hours) are measured:

Lifespans of Brand A bulbs (hours):

125001250011800118001320013200121001210012900129001190011900127001270012300123001310013100120001200012600126001220012200

Lifespans of Brand B bulbs (hours):

1210012100115001150012800128001190011900124001240011700117001200012000122001220011600116001230012300

The competitor performs a t-test at the 5%5\% significance level. It is assumed that the populations are normally distributed and have equal variances.

(f) Write down the null and alternative hypotheses for this test.

[2]
(g)

(g) Perform the t-test and state the conclusion, justifying your reasoning.

[4]

Question 5

MediumPaper 1 · calculator7 marks
(a)

A bank manager observes that the number of customers arriving at their main ATM during peak hours follows a Poisson distribution with a mean of 18 customers per hour.

To encourage more foot traffic, the bank launches a new promotional campaign. The manager wants to test if this campaign has increased the number of ATM arrivals. They decide to monitor the ATM for a single 4-hour peak period and use a 5% level of significance for their test.

(a) State the null and alternative hypotheses for the test.

[1]
(b)

(b) Find the probability that the bank manager will make a Type I error in their test conclusion.

[4]
(c)

During the 4-hour observation period, the bank manager recorded 85 customer arrivals at the ATM.

(c) State the bank manager's conclusion to the test. Justify your answer.

[2]

Question 6

HardPaper 2 · calculator21 marks
(a)

A company manufactures specialized medical devices. Each device consists of a main circuit board and a protective casing. The weight of the circuit board, CC, is normally distributed with a mean of 120120 g and a standard deviation of 55 g. The weight of the protective casing, PP, is normally distributed with a mean of 3030 g and a standard deviation of 22 g. The weights of the circuit board and the casing are independent.

Find the probability that a randomly chosen complete device has a total weight of less than 145145 g.

[5]
(b)

A batch of 1010 such devices is to be packed into a container. The container has a maximum weight capacity of 15101510 g. The weight of each device is independent.

Find the probability that the total weight of the 1010 devices is greater than the capacity of the container.

[4]
(c)(i)

The company sources a critical microchip component from two different suppliers, Supplier X and Supplier Y. An engineer claims that microchips from Supplier Y have a lower response time than those from Supplier X. To test this claim, a random sample is taken from each supplier.

The eight microchips in the sample from Supplier X have response times, in milliseconds (ms), of:

19.65,19.65,19.79,20.75,20.97,21.15,22.28,22.3719.65, 19.65, 19.79, 20.75, 20.97, 21.15, 22.28, 22.37.

Find the

(i) mean response time for the sample from Supplier X.

[3]
(c)(ii)

Find the

(ii) unbiased estimate of the population variance for the sample from Supplier X.

[3]
(d)

The seven microchips in the sample from Supplier Y have a mean response time of 19.519.5 ms and an unbiased estimate of the population standard deviation (sn−1s_{n-1}) of 1.051.05 ms.

Perform a suitable test, at the 5%5\% significance level, to test the engineer's claim that microchips from Supplier Y have a lower response time than those from Supplier X. You may assume the response times of microchips from each supplier are normally distributed with equal population variance.

[6]

Question 7

MediumPaper 1 · calculator6 marks
(a)

A materials scientist is developing a new synthetic fiber and claims its average tensile strength is 120 N. To test this claim, she takes a random sample of 25 fiber segments.

Given that the tensile strength of an individual fiber segment is xx newtons, the scientist found that, from her 25 samples, ∑x=2968\sum x = 2968 and ∑x2=352480\sum x^2 = 352480.

(a) Find an unbiased estimate for the mean tensile strength (μ\mu) of the fiber.

[1]
(b)

(b) Use the formula sn−12=∑x2−(∑x)2nn−1s_{n-1}^2 = \frac{\sum x^2 - \frac{(\sum x)^2}{n}}{n-1} to determine an unbiased estimate for the variance of the tensile strength of the fiber.

[2]
(c)

(c) Find a 95% confidence interval for μ\mu. You may assume that all conditions for a confidence interval have been met.

[2]
(d)

(d) Suggest, with justification, a valid conclusion that the materials scientist could make regarding her claim.

[1]

Question 8

HardPaper 2 · calculator18 marks
(a)

(a) The manager of "The Daily Grind" coffee shop suggested that the number of customers arriving at the shop during a 5-minute interval can be modelled by a Poisson distribution.

Suggest two observations that the manager may have made that led him to suggest this model.

[2]
(b)

(b) Now assume that the model is valid and that the mean number of customers arriving at the shop during a 5-minute interval is 1.51.5.

The manager observes customer arrivals during a 15-minute interval.

Calculate the probability that exactly 6 customers arrive during this 15-minute interval.

[3]
(c)

(c) Using the same model as in part (b), find the probability that fewer than 4 customers arrive during a 15-minute interval.

[2]
(d)

(d) Find the probability that in four consecutive 5-minute intervals, at least one customer arrives in each interval.

[3]
(e)

(e) Following a new marketing campaign, the manager wished to determine whether the mean number of customers arriving during a 5-minute interval had increased.

State the hypotheses for the test.

[2]
(f)

(f) Find the critical region for the test at the 5% significance level.

[3]
(g)

(g) Given that the mean number of customers per 5-minute interval has actually risen to 2.52.5, find the probability that the manager makes a Type II error.

[3]

Question 9

MediumPaper 1 · calculator7 marks
(a)

The number of incoming calls received by a tech support hotline in a particular 30-minute period is historically known to follow a Poisson distribution with a mean of 48 calls per hour.

A new automated system is implemented, and it is claimed that this system has decreased the number of calls requiring human intervention.

To test this claim, the number of calls, Y, requiring human intervention in a 30-minute period on a particular day will be recorded. The test will have the following hypotheses:

H0H_0: the mean number of calls requiring human intervention has not changed,

H1H_1: the mean number of calls requiring human intervention has decreased.

The alternative hypothesis will be accepted if Y≤17Y \leq 17.

Assuming the null hypothesis to be true, state the distribution of Y.

[1]
(b)

Find the probability of a Type I error.

[2]
(c)

Find the probability of a Type II error, if the number of calls now follows a Poisson distribution with a mean of 35 calls per hour.

[4]

Question 10

HardPaper 3 · calculator24 marks
(a)(i)

Ms. Anya Sharma, a school principal, wants to investigate if the number of hours students spend studying affects their exam scores. This question asks you to review Ms. Sharma's methods and conclusions.

Ms. Sharma obtained a list of students from her school. She contacted them and asked them to fill in an anonymous questionnaire. Participants were asked to state their weekly study hours and their most recent exam score (out of 100). Of the 250 students on the list, 11 replied.

Ms. Sharma's results are summarized in the following table:

Student IDWeekly Study Hours (X)Exam Score (Y)
1550
2765
3860
41078
51270
6655
7972
81180
9445
101385
111860

Describe one way in which Ms. Sharma could improve the reliability of her investigation.

[1]
(a)(ii)

Describe one criticism that can be made about the validity of Ms. Sharma's investigation.

[1]
(b)

Ms. Sharma classifies Student 11 as an outlier and removes their data from the analysis. Suggest one possible justification for her decision to remove it.

[1]
(c)(i)

For the remaining ten student responses in the table, Ms. Sharma calculates the mean exam score to be 6666. Calculate the mean weekly study hours for these remaining responses.

[2]
(c)(ii)

Determine the value of rr, Pearson's product-moment correlation coefficient, for these remaining responses.

[2]
(d)(i)

Ms. Sharma decides to carry out a hypothesis test on the correlation coefficient to investigate whether increased weekly study hours are associated with higher exam scores. State why the hypothesis test should be one-tailed.

[1]
(d)(ii)

State the null and alternative hypotheses for this test.

[2]
(d)(iii)

The critical value for this test, at the 5% significance level, is 0.549. Ms. Sharma assumes that the population is bivariate normal. Determine whether there is significant evidence of a positive correlation between weekly study hours and exam scores. Justify your answer.

[2]
(e)(i)

Ms. Sharma wants to create a model to predict how changing weekly study hours might affect exam scores. To do this, she assumes that weekly study hours, XX, is the independent variable and the exam score, YY, is the dependent variable.

She first considers a linear model of the form Y=aX+bY = aX + b. Use Ms. Sharma's data to find the value of aa and of bb.

[1]
(e)(ii)

Interpret, referring to study hours and exam scores, what the value of aa represents.

[1]
(e)(iii)

Ms. Sharma then considers a quadratic model of the form Y=cX2+dX+eY = cX^2 + dX + e. Find the value of cc, of dd and of ee.

[1]
(e)(iv)

Find the coefficient of determination for each of the two models she considers.

[2]
(e)(v)

Hence compare the two models.

[1]
(e)(vi)

Ms. Sharma decides to use the coefficient of determination to choose between these two models. Comment on the validity of her decision.

[1]
(f)(i)

After presenting the results of her investigation, a colleague questions whether Ms. Sharma's sample is representative of all students in the school. A report states that the mean weekly study hours for all students in the school is 99 hours. Ms. Sharma decides to carry out a test to determine whether her sample could realistically be taken from a population with a mean of 99 hours. State the name of the test which Ms. Sharma should use.

[1]
(f)(ii)

State the null and alternative hypotheses for this test.

[1]
(f)(iii)

Perform the test, using a 5% significance level, and state your conclusion in context.

[3]

Question 11

MediumPaper 1 · calculator6 marks
(a)

A factory produces two types of electronic components: standard (S) and premium (P). The weight of these components is a critical characteristic for quality control.

The weights of standard components are known to be normally distributed with a mean of 150 grams and a standard deviation of 5 grams.

The weights of premium components are known to be normally distributed with a mean of 165 grams and a standard deviation of 8 grams.

A quality control machine classifies a component as 'premium' if its weight is found to be above 158 grams; otherwise, it is classified as 'standard'.

The factory's quality control manager uses the null hypothesis that, in the absence of other information, a component is standard.

Calculate the probability of making a Type I error when classifying a component.

[2]
(b)

Calculate the probability of making a Type II error when classifying a component.

[2]
(c)

It is known that 80% of the components produced are standard, and 20% are premium.

Calculate the overall probability that a randomly selected component is misclassified by the machine.

[2]

Question 12

HardPaper 3 · calculator27 marks
(a)(i)

(a) Mr. Lee, the owner of "Sweet Delights" bakery, recorded the number of Mooncakes sold each day for a sample of 3030 days. The results are shown in the table below.

Number of Mooncakes soldFrequency
01
12
24
35
47
56
63
72

(a.i) Find the mean and variance for this sample data.

[2]
(a)(ii)

(a.ii) Hence, state why Mr. Lee might believe that the daily sales of Mooncakes follow a Poisson distribution.

[1]
(b)

(b) State one assumption that Mr. Lee needs to make about the sales of Mooncakes to support his belief that it follows a Poisson distribution.

[1]
(c)

(c) Mr. Lee knows from his historic sales records that the bakery sells an average of 3.93.9 Mooncakes each day. The following table shows the expected frequency of Mooncakes sold each day during a 100100-day period, assuming a Poisson distribution with mean 3.93.9.

Number of Mooncakes sold<11234567≥8\ge 8
Expected frequencya7.8947.89415.39415.39420.01220.012b15.21915.2199.8939.8935.5125.512c

Find the value of a, of b, and of c. Give your answers to 3 decimal places.

[5]
(d)(i)

(d) Mr. Lee decides to carry out a χ2\chi^2 goodness of fit test at the 5%5\% significance level to see whether the daily sales of Mooncakes follow a Poisson distribution with mean 3.93.9. He collects observed frequencies for 100100 days, which are given in the table below.

Number of Mooncakes sold<223456≥7\ge 7
Observed frequency12132317121013
Expected frequency9.9189.91815.39415.39420.01220.01219.51219.51215.21915.2199.8939.89310.05210.052

(d.i) Write down the number of degrees of freedom for his test.

[1]
(d)(ii)

(d.ii) Perform the χ2\chi^2 goodness of fit test and state, with reason, a conclusion.

[7]
(e)(i)

(e) Mr. Lee claims that a new social media advertising campaign, costing 250250 THB per day, will increase the number of Mooncakes sold. However, his business partner, Ms. Chen, claims that the advertising will not increase the bakery's overall profit.

Ms. Chen agrees to run the campaign for the next 4040 days. During that time, Mr. Lee records that the bakery sells a total of 180180 Mooncakes, with a profit of 4545 THB on each Mooncake sold.

Mr. Lee wants to carry out an appropriate hypothesis test to determine whether the number of Mooncakes sold during the 4040 days increased when compared with the historic sales records (mean 3.93.9 Mooncakes per day).

By finding a critical value, perform this test at a 5%5\% significance level.

[6]
(f)

(f) Hence state the probability of a Type I error for this test.

[1]
(g)

(g) By considering the claims of both Mr. Lee and Ms. Chen, explain whether the advertising campaign was beneficial to the bakery.

[3]

Question 13

MediumPaper 1 · calculator8 marks
(a)

A manufacturing company produces electronic components. Historically, the defect rate for a specific component has been 15%. A new production line is implemented, and the quality control manager wants to test if the defect rate has decreased. She assumes that the defect status of each component is independent of others.

(a) Write down suitable hypotheses for this test.

[2]
(b)

(b) The quality control manager decides to take a random sample of 120 components. She will reject the null hypothesis if fewer than 12 components are found to be defective.

Find the probability that she makes a Type I error.

[3]
(c)

(c) In fact, the new production line successfully reduced the defect rate to 10%.

Find the probability that she makes a Type II error.

[3]

Question 14

HardPaper 3 · calculator28 marks
(a)(i)

(a) TechInnovate is considering collecting more data for their analysis.

(i) State one advantage of increasing the sample size.

[1]
(a)(ii)

(ii) State one disadvantage of increasing the sample size.

[1]
(b)

(b) The production manager at Plant Alpha recorded the time, in minutes, taken to produce a batch of electronic components for 10 randomly selected batches:

18.2,19.5,17.8,20.1,18.5,19.0,17.5,20.5,18.8,19.318.2, 19.5, 17.8, 20.1, 18.5, 19.0, 17.5, 20.5, 18.8, 19.3

Find the value of sn−1s_{n-1} for this sample from Plant Alpha.

[2]
(c)

(c) A manager claims that Plant Alpha's production times are more consistent than Plant Beta's. Given that the sample standard deviation (sn−1s_{n-1}) for Plant Beta's production times is 1.051.05 minutes, make one criticism of this claim.

[1]
(d)(i)

(d) TechInnovate wants to compare the mean production times of Plant Alpha and Plant Beta using a pooled t-test.

(i) State the condition regarding population variances required to use a pooled t-test.

[1]
(d)(ii)

(ii) Given that for Plant Beta, a sample of 1212 batches yielded a mean production time of xˉB=19.3\bar{x}_B = 19.3 minutes and a sample standard deviation of sB=1.05s_B = 1.05 minutes, state whether TechInnovate should use a pooled t-test in this case. Justify your answer.

[2]
(e)(i)

(e) TechInnovate believes Plant Alpha has a lower mean production time than Plant Beta.

(i) State appropriate null and alternative hypotheses for the pooled t-test.

[2]
(e)(ii)

(ii) Find the p-value.

[2]
(e)(iii)

(iii) Given that the test is carried out at the 5% significance level, state the appropriate conclusion in context. Justify your answer.

[2]
(f)(i)

(f) The company also investigates the relationship between operator experience (in years) and the number of defective items produced per day. A sample of 8 operators yielded the following data:

Operator Experience (years)Number of Defective Items
215
510
313
87
118
69
412
78

(i) Assuming all requirements are met, perform a test at the 5% significance level to determine if there is a linear correlation between operator experience and the number of defective items. State the hypotheses and justify your conclusion.

[4]
(f)(ii)

(ii) If the requirements for this test are not met, state an alternative test that could be used.

[1]
(g)

(g) For the data in (f.i), the equation of the least squares regression line of defective items (DD) on operator experience (EE) is D=−1.5E+18.25D = -1.5E + 18.25. Give, in context, an interpretation of the gradient −1.5-1.5 in this model.

[1]
(h)(i)

(h) TechInnovate uses a baseline model to predict the number of defective items (DpredD_{pred}) for a batch based on its size (SS): Dpred=0.5S+10D_{pred} = 0.5S + 10. The "Quality Deviation" (QQ) for a batch is defined as Q=Dpred−DactualQ = D_{pred} - D_{actual}. A positive Quality Deviation indicates better-than-expected quality.

(i) Show that for a batch of 150150 components from Plant Beta that produced 8080 defective items, the Quality Deviation is 5.05.0.

[2]
(h)(ii)

(ii) To compare quality control, samples of Quality Deviation scores were collected:

  • Plant Alpha: nQA=15n_{QA} = 15, xˉQA=4.5\bar{x}_{QA} = 4.5, sQA=1.2s_{QA} = 1.2
  • Plant Beta: nQB=18n_{QB} = 18, xˉQB=3.8\bar{x}_{QB} = 3.8, sQB=1.1s_{QB} = 1.1

Assuming that the appropriate requirements are met, use a pooled t-test at a 5% significance level to determine if the mean Quality Deviation is higher in Plant Alpha than in Plant Beta. Write down your null and alternative hypotheses and justify your conclusion.

[4]
(i)

(i) Using the results from parts (e) and (h.ii), explain how each plant could claim they are performing better than the other plant.

[2]

Question 15

MediumPaper 1 · calculator6 marks
(a)

Ms. Chen, a small business owner, wants to investigate if there is a monotonic relationship between the monthly advertising budget and the number of units of a new product sold. She collects data for eight months, as shown in Table 1.

Table 1: Monthly Advertising Budget and Units Sold

MonthAdvertising Budget (in $100s)Units Sold
12.5120
23.0150
31.895
44.2170
53.5110
62.0180
74.8230
83.2130

Ms. Chen decides to calculate the Spearman's rank correlation coefficient. Complete the table of ranks shown in Table 2.

Table 2: Ranks for Advertising Budget and Units Sold

MonthRank of Advertising Budget (RXR_X)Rank of Units Sold (RYR_Y)
133
25
311
47
52
62
788
84
[1]
(b)

Calculate the value of rsr_s, Spearman's rank correlation coefficient.

[2]
(c)

Ms. Chen believes that a higher advertising budget leads to more units sold. She carries out a hypothesis test using a 10% significance level with the following null hypothesis:

H0H_0: In the population, there is no monotonic relationship between the monthly advertising budget and the number of units sold.

Write down Ms. Chen's alternative hypothesis.

[1]
(d)

The critical value of rcr_c for this test is 0.643.

State the conclusion of the hypothesis test, giving a reason.

[2]

Question 16

HardPaper 2 · calculator18 marks
(a)

The battery life, in hours, of a particular smartphone model, LL, can be modelled by a normal distribution with a mean of 24 hours and a standard deviation of 2 hours.

(a) Find the probability that a randomly selected smartphone has a battery life greater than 27 hours.

[2]
(b)(i)

Two smartphones are selected at random and independently of each other.

(b) (i) Find the probability that both smartphones have a battery life greater than 27 hours.

[2]
(b)(ii)

(b) (ii) Find the probability that their total battery life is greater than 52 hours.

[4]
(c)

A software update is released which is claimed to improve battery life. The manufacturer decides to take a random sample of 20 smartphones to test this claim at the 1% significance level, assuming the standard deviation of the battery life has not changed.

(c) Write down the null and alternative hypotheses for the test.

[1]
(d)

(d) Find the critical region for this test.

[4]
(e)

Unknown to the manufacturer, the software update has resulted in all smartphones having a 5% longer battery life than the original model.

(e) Find the mean and standard deviation of the battery life for smartphones with the update.

[3]
(f)

(f) Find the probability of a Type II error in the manufacturer’s test.

[2]

Question 17

MediumPaper 1 · calculator6 marks
(a)

A botanist is investigating the effectiveness of two different soil compositions, Soil A and Soil B, on the growth of a particular plant species. They hypothesize that Soil A will lead to a greater mean plant height after four weeks compared to Soil B. The botanist grows a random sample of plants in each soil type and measures their heights (in cm) after four weeks. The results are shown in the table below.

Soil TypePlant Heights (cm)
Soil A25.5, 23.9, 26.6, 27.2, 25.1, 24.5, 25.4, 25.9, 24.8, 26.1
Soil B22.3, 23.4, 22.8, 23.2, 23.7, 24.1, 22.9, 23.5, 23.1, 22.5, 23.8, 23.0

The botanist performs a one-tailed t-test at a 5% level of significance. It is assumed that the plant heights are normally distributed and the samples have equal variances.

State the null and alternative hypotheses.

[2]
(b)

Calculate the p-value for this test.

[2]
(c)

State the conclusion of the test in the context of the question. Justify your answer.

[2]

Question 18

MediumPaper 1 · calculator7 marks

A company, "Electro-Tech", manufactures electronic components. Their quality control department claims that the average number of defective components in a standard batch of 1000 units follows a Poisson distribution with a mean of 1515 defects per batch. A new supplier's components are being tested. To assess their quality, 88 standard batches of components from the new supplier are randomly selected and inspected. A total of 135135 defective components are found across these 88 batches.

Test the claim that the new supplier's components have an average of 1515 defects per batch against the suspicion that they have more defects, at the 5%5\% significance level. In your answer, you should include whether this is a one-tailed or two-tailed test, the test hypotheses, calculation of the pp-value, and the conclusion (with a reason) of the test.

Question 19

MediumPaper 2 · calculator13 marks
(a)(i)

The battery life, in hours, of eight Brand A batteries was recorded as 45.2,47.8,48.5,49.1,50.3,51.7,52.4,54.045.2, 47.8, 48.5, 49.1, 50.3, 51.7, 52.4, 54.0.

The battery life, in hours, of 1010 Brand B batteries was recorded as 43.5,44.9,46.1,47.2,48.0,48.8,49.5,50.1,51.0,52.343.5, 44.9, 46.1, 47.2, 48.0, 48.8, 49.5, 50.1, 51.0, 52.3.

(a) Find the sample mean battery life for

(i) Brand A batteries.

[2]
(a)(ii)

(ii) Brand B batteries.

You can assume that both sets of values have a common unknown variance.

[2]
(b)

(b) Carry out a test at the 10%10\% significance level to determine if the population mean battery life of Brand A is longer than that of Brand B. In your answer, you should include the test used (with a reason), the test hypotheses, the pp-value and the conclusion of the test (with a reason).

[7]
(c)

(c) State if your conclusion would have been any different if working at the 5%5\% significance level.

[2]

Question 20

MediumPaper 1 · calculator5 marks

A global tech company has historically received customer complaints for its flagship software product following a Poisson distribution with a mean of 99 per week. Recently, after a major software update, the company's quality assurance team wants to investigate if the number of complaints has increased.

Over a period of 33 weeks following the update, they recorded a total of 3535 complaints.

Test at the 5%5\% significance level the hypothesis that the mean number of complaints has increased.

3 more Critical Values/Regions, Population Mean Tests (normal/poisson), Proportion Tests (binomial), Correlation Hypothesis Testing, and Type I/II Errors questions in the app

Every answer is marked mark by mark, IB-style, and the AI tutor helps when you are stuck.

Where marks are lost

  • Rounding an intermediate value and then using it in a later part. Costs a mark every time, and AI's multi-part modelling questions give it more chances to happen than AA's shorter, more self-contained ones.
  • Using your own wrong value after failing a "show that." All follow through is withdrawn for the rest of that question.
  • Leaving an answer in calculator notation. Never accepted in a final answer, and AI's constant calculator use makes this the easiest slip in the whole subject.
Free. Every IB subject.
No card, no trial that runs out. Just a free account.
  • 50 marked answers a month
    Marked mark by mark, IB-style
  • Hints and mark schemes
    On every part of every question
  • 3,000+ questions
    All 6 subjects, SL and HL, mapped to the syllabus
  • Progress that adapts
    Your Study Profile picks what to practise next

Practise this topic as a session

Pick a difficulty and paper, and FourtyFive tracks your progress on this topic as you go.

or with email
FAQ

Questions,
answered.

Can't find what you're looking for? Email our student team.

What does Critical Values/Regions, Population Mean Tests (normal/poisson), Proportion Tests (binomial), Correlation Hypothesis Testing, and Type I/II Errors cover in IB Maths AI?

Hypothesis Test: Uses sample data to evaluate a statement about a population parameter. Null Hypothesis (H_0): Default statement (no change, no difference, or no correlation). Alternative Hypothesis (H_1): Opposes H_0 (>, <, or ≠).

Is Critical Values/Regions, Population Mean Tests (normal/poisson), Proportion Tests (binomial), Correlation Hypothesis Testing, and Type I/II Errors SL or HL?

Critical Values/Regions, Population Mean Tests (normal/poisson), Proportion Tests (binomial), Correlation Hypothesis Testing, and Type I/II Errors is HL only. SL students are not examined on it.

How do I revise Critical Values/Regions, Population Mean Tests (normal/poisson), Proportion Tests (binomial), Correlation Hypothesis Testing, and Type I/II Errors for IB Maths AI?

Start from the core idea: hypothesis Test: Uses sample data to evaluate a statement about a population parameter. In the exam: the hardest statistics content in the course, and a Paper 3 favourite. Critical regions for discrete distributions are the fiddly part: the region is built up one value at a time until adding another would push the Type I error above the significance level, which is why the guide spells out the maximising rule. Then practise exam-style questions, easiest first, writing out every step of your working before you check it.

How does FourtyFive help me practise Critical Values/Regions, Population Mean Tests (normal/poisson), Proportion Tests (binomial), Correlation Hypothesis Testing, and Type I/II Errors?

FourtyFive has 23 Critical Values/Regions, Population Mean Tests (normal/poisson), Proportion Tests (binomial), Correlation Hypothesis Testing, and Type I/II Errors questions. Every answer you write is marked mark by mark, IB-style, and you see where each mark was won or lost. Every part has a hint, the AI tutor helps you through the step you are stuck on, and your Study Profile picks what to practise next.

Is FourtyFive free for Critical Values/Regions, Population Mean Tests (normal/poisson), Proportion Tests (binomial), Correlation Hypothesis Testing, and Type I/II Errors practice?

Yes. A free account gives you 50 marked answers a month, and you do not need a card to sign up.

Can I handwrite Critical Values/Regions, Population Mean Tests (normal/poisson), Proportion Tests (binomial), Correlation Hypothesis Testing, and Type I/II Errors answers on an iPad?

Yes. In the FourtyFive iPad app you write your working by hand with Apple Pencil, the way you would on paper, and it is marked the same way.

Start with the IB question
bank built for you.

Free to start, no card needed. Thousands of syllabus-mapped questions, AI Examiner marking, your weakest topics first.