Skip to content
  1. IB Question Bank
  2. Maths AI
  3. Statistics & Probability
Topic 4.01 · SL and HL

Stats basics (population, 5 sampling techniques, outlier definition): notes and practice questions

Summary
  • Qualitative data is descriptive, non-numerical.
  • Samples may not be representative of the entire population; minimise bias with random sampling.
  • Stratified sampling divides the population into disjoint groups, then samples each group such that:
  • The proportion sampled from a group equals the proportion of the population in that group.
  • Quartiles divide data into four equal sections:
  • Q1Q_1: Splits the lowest 25%.
  • Q2Q_2 (Median): Splits the lowest 50%.
  • Q3Q_3: Splits the lowest 75%.
  • The Interquartile Range (IQR) measures the spread of the middle 50% of data:

IQR=Q3−Q1IQR = Q_3 - Q_1

  • A value is an outlier if it falls outside these boundaries:
  • Lower Boundary: Q1−1.5×IQRQ_1 - 1.5 \times IQR
  • Upper Boundary: Q3+1.5×IQRQ_3 + 1.5 \times IQR
  • Box plots display minimum, Q1Q_1, median, Q3Q_3, and maximum values.
  • Outliers are explicitly marked with a cross (×\times).
  • Whiskers extend to the next smallest/largest value in the data set after any outliers are identified.
  • Use a Graphical Display Calculator (GDC) to find quartiles and check outliers via its box plot function.
  • A single sample might not perfectly represent the whole population.
  • When drawing box plots with outliers, whiskers must stop at the next valid data point, not the outlier itself.

How it is examined

Cheap marks and easy marks to lose. Naming a sampling technique is one mark, and saying why it is or is not appropriate here is another, and the second needs a reference to the actual context. The outlier rule is applied numerically, so a question can ask a student to test a specific value against Q1−1.5 IQRQ_1 - 1.5\,\text{IQR} and Q3+1.5 IQRQ_3 + 1.5\,\text{IQR}. Take care with the word "random": simple random is one of the five named methods, not a synonym for unbiased.

Key ideas
  • The concepts of population, sample, random sample, and discrete and continuous data.
  • The reliability of data sources and bias in sampling.
  • The interpretation of outliers.
  • Sampling techniques and their effectiveness.

Linking questions

  • Links to other subjects: descriptive statistics and random samples (biology, psychology, sports exercise and health science, environmental systems and societies, geography, economics, business management); research methodologies (psychology).
  • Aim 8: misleading statistics, and problems caused by unrepresentative samples, for example the Google flu predictor, the 1936 US presidential election, the Literary Digest against George Gallup, the Boston "pot-hole" app.
  • International-mindedness: the Kinsey report and its sampling techniques.
  • TOK: why have mathematics and statistics sometimes been treated as separate subjects? How easy is it to be misled by statistics? Is it ever justifiable to use statistics to mislead deliberately?

Practice questions

23 questions · 15 medium · 8 hard
Showing 20 of 20

Question 1

MediumPaper 1 · calculator8 marks
(a)

A renowned artisanal bakery claims that only 5% of its specialty sourdough loaves have minor cosmetic imperfections (e.g., slight cracks, uneven browning). A local restaurant owner, who regularly purchases these loaves, decides to test this claim. For their latest delivery, the owner inspects a batch of 150 loaves and finds 12 loaves with cosmetic imperfections.

(a) Identify the type of sampling used by the restaurant owner.

[1]
(b)

(b) State the null and alternative hypotheses for this test.

[2]
(c)

(c) Calculate the p-value for this hypothesis test, assuming cosmetic imperfections occur independently.

[3]
(d)

(d) The restaurant owner performs the test at the 5% significance level. State the conclusion of the test, giving a reason.

[2]

Question 2

HardPaper 2 · calculator22 marks
(a)

A logistics company recorded the delivery times (in minutes) for a large batch of packages. The data is grouped in the frequency table below:

Delivery Time (minutes) | Frequency

---|---

20≤t<2520 \le t < 25 | 88

25≤t<3025 \le t < 30 | 2222

30≤t<3530 \le t < 35 | 5555

35≤t<4035 \le t < 40 | 9898

40≤t<4540 \le t < 45 | 135135

45≤t<5045 \le t < 50 | 110110

50≤t<5550 \le t < 55 | 7070

55≤t<6055 \le t < 60 | 3535

60≤t<6560 \le t < 65 | 1515

65≤t<7065 \le t < 70 | 22

(a) Calculate estimates of the mean and standard deviation of the delivery times.

[4]
(b)

(b) Construct a cumulative frequency table for the data, and use it to draw a cumulative frequency curve.

Cumulative frequency curve placeholder. X-axis: Delivery Time (minutes), Y-axis: Cumulative Frequency.
[4]
(c)(i)

(c) Use your graph to estimate:

(i) the median delivery time

[2]
(c)(ii)

(ii) the lower and upper quartile of the delivery times

[2]
(c)(iii)

(iii) the interquartile range

[1]
(c)(iv)

(iv) the 8585th percentile of delivery times.

[2]
(d)

(d) Draw a box-and-whisker plot of the data.

Box and whisker plot placeholder. X-axis: Delivery Time (minutes).
[3]
(e)

(e) Determine, with reasons, whether any customers could be considered outliers.

[4]

Question 3

MediumPaper 1 · calculator19 marks
(a)(i)

[Maximum mark: 19]

A tech company recorded the time (in minutes) 180 customers spent completing a new online feedback survey. The data was compiled into the following cumulative frequency graph.

Cumulative frequency graph showing time in minutes on x-axis and cumulative frequency on y-axis (from 0 to 180)

(a) Use the graph to find

(i) the median time;

[4]
(a)(ii)

(ii) the lower quartile;

[4]
(a)(iii)

(iii) the upper quartile;

[4]
(a)(iv)

(iv) the interquartile range.

[4]
(b)

Sarah completed the survey in 1.5 minutes.

(b) Determine whether Sarah's time is an outlier.

[3]

Question 4

HardPaper 2 · calculator21 marks
(a)

Dr. Anya Sharma, a sports scientist, is investigating the relationship between training habits and performance in junior athletes. She wants to collect data on the weekly training hours of junior swimmers. She decides to interview every 5th swimmer entering the training facility until she has a sample of 50 swimmers.

State the sampling method Dr. Sharma has used.

[1]
(b)

Dr. Sharma constructed the following box and whisker diagram to show the weekly training hours (in hours) of a sample of junior swimmers.

A box and whisker diagram showing weekly training hours. The minimum is 2, the first quartile (Q1) is 4, the median is 6, the third quartile (Q3) is 9, and the maximum is 12.

Write down the median weekly training hours.

[1]
(c)

Calculate the interquartile range for the weekly training hours.

[2]
(d)

One swimmer in the sample reported training for 15 hours per week. Dr. Sharma believes this swimmer's training time is not an outlier.

Determine whether Dr. Sharma is correct. Support your reasoning.

[4]
(e)

Dr. Sharma also collected data on the average weekly training hours (xx) and the competition score (yy) for a group of athletes. These data are represented on the scatter diagram.

A scatter diagram showing competition score (y-axis from 0 to 120) versus weekly training hours (x-axis from 0 to 25). The points show a general negative correlation, with data points roughly between 5 and 20 hours.

Describe the correlation between weekly training hours and competition score.

[1]
(f)

Dr. Sharma correctly calculates the equation of the regression line yy on xx for these athletes to be y=−2.5x+110y = -2.5x + 110. She uses the equation to estimate the competition score for an athlete who trains 3 hours per week.

Find the competition score calculated by Dr. Sharma.

[2]
(g)

State whether it is valid to use the regression line yy on xx for Dr. Sharma's estimate in part (f). Give a reason for your answer, assuming the original data for training hours ranged from 5 to 20 hours.

[2]
(h)

Dr. Sharma investigated the relationship between an athlete's national competition rank and their average daily protein intake (in grams). She collected data for eight athletes, as shown in the table.

AthleteABCDEFGH
Competition Rank (RcompR_{comp})12345678
Protein Intake (g) (PintakeP_{intake})180150200160140190170130

Dr. Sharma intends to analyse the data using Spearman's rank correlation coefficient, rsr_s.

Copy and complete the information in the following table.

AthleteABCDEFGH
Rank - Competition Rank1
Rank - Protein Intake
[2]
(i)(i)

Calculate the value of rsr_s.

[3]
(i)(ii)

Interpret your result.

[3]

Question 5

MediumPaper 2 · calculator15 marks
(a)

A fitness enthusiast, Alex, recorded the number of steps (in hundreds) he took each day over a period of 2828 days. The data is given below in ascending order:

55,60,62,65,68,70,72,75,78,80,82,85,88,90,92,95,98,100,102,105,110,115,120,125,130,140,160,25055, 60, 62, 65, 68, 70, 72, 75, 78, 80, 82, 85, 88, 90, 92, 95, 98, 100, 102, 105, 110, 115, 120, 125, 130, 140, 160, 250

(a) Find the median number of steps.

[2]
(b)

(b) Find the lower quartile (Q1).

[2]
(c)

(c) Find the upper quartile (Q3).

[2]
(d)

(d) Find the range of the data.

[2]
(e)

(e) Determine whether there are any outliers in the data.

[4]
(f)

(f) Draw a box-and-whisker diagram for the above data, marking any outliers as required.

[3]

Question 6

HardPaper 2 · calculator19 marks
(a)

A manufacturing company inspects the first 50 items produced each morning for defects.

(a) State the sampling method being used.

[1]
(b)(i)

The company uses an automated machine to test products for defects. This machine is not perfect.

It is known that 3% of all products manufactured are defective (D).

If a product is defective, the machine correctly identifies it as defective (tests positive, T+T+) 98% of the time.

If a product is not defective (D'), the machine incorrectly identifies it as defective (tests positive, T+T+) 1% of the time.

The tree diagram shows some of this information.

Tree diagram showing probabilities of product defect and test results

(b) (i) Write down the value of P(D′)P(D').

[1]
(b)(ii)

(ii) Write down the value of P(T−∣D)P(T-|D).

[1]
(b)(iii)

(iii) Write down the value of P(T+∣D′)P(T+|D').

[1]
(b)(iv)

(iv) Write down the value of P(T−∣D′)P(T-|D').

[1]
(c)(i)

(c) Use the tree diagram to find the probability that a randomly selected product:

(i) is not defective and tests positive.

[2]
(c)(ii)

(ii) tests negative.

[3]
(c)(iii)

(iii) is defective given that it tested negative.

[3]
(d)

(d) The company finds the actual number of defective products in their sample is different than predicted by the tree diagram. Explain why this might be the case.

[1]
(e)

(e) The factory manager surveyed all employees on a particular shift. All employees on this shift worked in at least one of these departments: Assembly (A), Painting (P), or Quality Control (Q). It was found that:

  • 85 employees worked in Assembly;
  • 55 employees worked in Painting;
  • 35 employees worked in Quality Control;
  • 10 employees worked in all three departments;
  • 20 employees worked in Assembly and Quality Control but not Painting;
  • 15 employees worked in Assembly and Painting but not Quality Control;
  • 4 employees worked only in Quality Control.

Draw a Venn diagram to illustrate this information, placing all relevant information on the diagram.

[3]
(f)

(f) Find the total number of employees on this shift.

[2]

Question 7

MediumPaper 1 · calculator12 marks
(a)

(a) State whether the following set of data is discrete or continuous, and, in each case, construct a frequency table.

The number of cars passing a specific checkpoint in 10-minute intervals during rush hour:

555667799910111112121212125 \quad 5 \quad 5 \quad 6 \quad 6 \quad 7 \quad 7 \quad 9 \quad 9 \quad 9 \quad 10 \quad 11 \quad 11 \quad 12 \quad 12 \quad 12 \quad 12 \quad 12

[4]
(b)

(b) State whether the following set of data is discrete or continuous, and, in each case, construct a frequency table using appropriate class intervals.

The heights of saplings (in cm) in a plant nursery:

15.717.318.419.120.221.322.624.425.226.026.527.528.529.029.715.7 \quad 17.3 \quad 18.4 \quad 19.1 \quad 20.2 \quad 21.3 \quad 22.6 \quad 24.4 \quad 25.2 \quad 26.0 \quad 26.5 \quad 27.5 \quad 28.5 \quad 29.0 \quad 29.7

[4]
(c)

(c) State whether the following set of data is discrete or continuous, and, in each case, construct a frequency table.

The number of correct answers on a 10-question multiple-choice quiz for a group of students:

4444555555667889999104 \quad 4 \quad 4 \quad 4 \quad 5 \quad 5 \quad 5 \quad 5 \quad 5 \quad 5 \quad 6 \quad 6 \quad 7 \quad 8 \quad 8 \quad 9 \quad 9 \quad 9 \quad 9 \quad 10

[4]

Question 8

HardPaper 2 · calculator15 marks
(a)

A quality control manager at a manufacturing plant wants to assess the consistency of a new batch of electronic components. He decides to test ten components, ensuring that five are selected from Production Line A and five from Production Line B. The manager instructs the supervisors of each line to provide the required number of components from their current production.

(a) Name the type of sampling that best describes the method used by the quality control manager.

[1]
(b)(i)

The weights, in grams, of the ten components selected for the test are:

148,153,161,155,142,160,150,157,145,159148, 153, 161, 155, 142, 160, 150, 157, 145, 159.

(b) For these ten components, find

(i) the mean weight.

[2]
(b)(ii)

(ii) the standard deviation of the weights.

[2]
(c)

The target weight for these components is 155155 g. The manager is concerned that the components might be consistently underweight. Perform an appropriate test at the 10%10\% significance level to see if the mean weight of the components produced is less than the target weight. It can be assumed that the weights come from a normal population.

[5]
(d)

State one reason why the test performed in part (c) might not be valid.

[1]
(e)(i)

Two additional components are tested at a later date. The mean weight for all twelve components is 154.5154.5 g and the standard deviation is 7.27.2 g.

For further analysis, a 'quality score' for the twelve components is obtained by multiplying the weights by 1.51.5 and subtracting 5050.

(e) For the twelve components, find

(i) their mean quality score.

[2]
(e)(ii)

(ii) the standard deviation of their quality score.

[2]

Question 9

MediumPaper 2 · calculator10 marks
(a)

The scores (out of 100) for 28 students on a recent mathematics test were recorded as follows:

55,60,62,65,65,68,70,70,72,73,75,75,76,78,79,80,80,82,83,85,85,86,88,90,92,95,96,9855, 60, 62, 65, 65, 68, 70, 70, 72, 73, 75, 75, 76, 78, 79, 80, 80, 82, 83, 85, 85, 86, 88, 90, 92, 95, 96, 98

(a) Find the mean score of the students.

[2]
(b)

(b) Find the median score of the students.

[2]
(c)

(c) Find the interquartile range.

[4]
(d)

(d) The score of another student was added to the data but found to be an outlier.

Find the least possible score this new student could have, given that it is higher than each of the others.

[2]

Question 10

HardPaper 2 · calculator21 marks
(a)(i)

The lifespans, tt, of 250 LED light bulbs are recorded in the following table.

Lifespan (hours)Frequency
0≤t<10000 \le t < 100020
1000≤t<15001000 \le t < 150060
1500≤t<20001500 \le t < 200090
2000≤t<25002000 \le t < 250055
2500≤t<30002500 \le t < 300025

This table is used to create a cumulative frequency graph.

Write down the mid-interval value of the class 0≤t<10000 \le t < 1000.

[1]
(a)(ii)

Calculate an estimate of the mean lifespan of the 250 light bulbs.

[3]
(b)

Use the cumulative frequency curve (which would be provided in an exam) to estimate the interquartile range. Assume the lower quartile (Q1Q_1) is 13001300 hours and the upper quartile (Q3Q_3) is 21502150 hours.

[3]
(c)

A light bulb from the data set had a lifespan of 34003400 hours.

Use your answer to part (b) to estimate whether this light bulb's lifespan is an outlier for this data. Justify your answer.

[3]
(d)

It is believed that the lifespans of these LED light bulbs follow a normal distribution with mean 17401740 hours and standard deviation 450450 hours.

It is decided to perform a χ2\chi^2 goodness of fit test on the data to determine whether this sample of 250 light bulbs could have plausibly been drawn from an underlying distribution N(1740,4502)N(1740, 450^2).

Write down the null and the alternative hypotheses for the test.

[2]
(e)(i)

As part of the test, the following table is created.

Lifespan of light bulb (hours)Observed frequencyExpected frequency
t<1000t < 10002014.0
1000≤t<15001000 \le t < 15006060.1
1500≤t<20001500 \le t < 200090a
2000≤t<25002000 \le t < 25005560.1
t≥2500t \ge 250025b

Find the value of aa and the value of bb. Give your answers to one decimal place.

[5]
(e)(ii)

Hence, perform the test to a 5% significance level, clearly stating the conclusion in context.

[4]

Question 11

MediumPaper 1 · calculator8 marks
(a)

P1: The box-and-whisker diagram below illustrates the daily screen time (in hours) for a sample of teenagers.

Box-and-whisker diagram showing values 2, 3.5, 5, 7, 9.5 with an x-axis from 0 to 14 in increments of 2

(a) Find the range of daily screen times.

[2]
(b)

(b) Find the interquartile range (IQR) of daily screen times.

[2]
(c)

(c) Find the percentage of teenagers who spend between 3.53.5 and 77 hours on screen time daily.

[1]
(d)

(d) A new survey participant reported spending 1212 hours on screen time daily.

Determine whether this time would be counted as an outlier.

[3]

Question 12

HardPaper 2 · calculator21 marks
(a)(i)

(a) The scores, ss, of 200 students on a mathematics test are recorded in the following table.

Score (ss)Frequency
20≤s<4020 \le s < 4015
40≤s<6040 \le s < 6035
60≤s<8060 \le s < 8060
80≤s<10080 \le s < 10050
100≤s<120100 \le s < 12030
120≤s<140120 \le s < 14010

(i) Write down the mid-interval value of 60≤s<8060 \le s < 80.

[3]
(a)(ii)

(ii) Calculate an estimate of the mean score of the 200 students.

[3]
(b)

(b) The data from this table is used to create a cumulative frequency graph. From this graph, the first quartile (Q1Q_1) is estimated to be 6060 and the third quartile (Q3Q_3) is estimated to be 9696.

Use these values to estimate the interquartile range (IQR).

[2]
(c)

(c) A student, Elara, scored 155155 on the test.

Use your answer to part (b) to estimate whether Elara's score is an outlier for this data. Justify your answer.

[3]
(d)

(d) It is believed that the scores of students on this mathematics test follow a normal distribution with mean 77.577.5 and standard deviation 2020.

It is decided to perform a χ2\chi^2 goodness of fit test on the data to determine whether this sample of 200 students could have plausibly been drawn from an underlying distribution N(77.5,202)N(77.5, 20^2).

Write down the null and the alternative hypotheses for the test.

[2]
(e)

(e) As part of the test, the following table is created, where some categories have been combined to ensure expected frequencies are not too low.

Score (ss)Observed FrequencyExpected Frequency
s<40s < 40156.08
40≤s<6040 \le s < 603532.08
60≤s<8060 \le s < 8060a
80≤s<10080 \le s < 1005063.99
s≥100s \ge 10040b

(i) Find the value of aa and the value of bb.

(ii) Hence, perform the test to a 5% significance level, clearly stating the conclusion in context.

[8]

Question 13

MediumPaper 1 · calculator10 marks
(a)

Define, as fully as you can, the terms random sampling, stratified sampling, and systematic sampling.

[5]
(b)(i)

A large university wants to conduct a survey to understand student satisfaction with the campus library services. The university initially considers using systematic sampling by selecting every 50th student from an alphabetical list of all enrolled students.

Suggest one reason why this method might not be appropriate for surveying student satisfaction with library services.

[2]
(b)(ii)

The research team is now considering either random sampling or stratified sampling. Determine which of these two methods would be more appropriate for this investigation. Justify your answer.

[3]

Question 14

HardPaper 3 · calculator24 marks
(a)(i)

Ms. Anya Sharma, a school principal, wants to investigate if the number of hours students spend studying affects their exam scores. This question asks you to review Ms. Sharma's methods and conclusions.

Ms. Sharma obtained a list of students from her school. She contacted them and asked them to fill in an anonymous questionnaire. Participants were asked to state their weekly study hours and their most recent exam score (out of 100). Of the 250 students on the list, 11 replied.

Ms. Sharma's results are summarized in the following table:

Student IDWeekly Study Hours (X)Exam Score (Y)
1550
2765
3860
41078
51270
6655
7972
81180
9445
101385
111860

Describe one way in which Ms. Sharma could improve the reliability of her investigation.

[1]
(a)(ii)

Describe one criticism that can be made about the validity of Ms. Sharma's investigation.

[1]
(b)

Ms. Sharma classifies Student 11 as an outlier and removes their data from the analysis. Suggest one possible justification for her decision to remove it.

[1]
(c)(i)

For the remaining ten student responses in the table, Ms. Sharma calculates the mean exam score to be 6666. Calculate the mean weekly study hours for these remaining responses.

[2]
(c)(ii)

Determine the value of rr, Pearson's product-moment correlation coefficient, for these remaining responses.

[2]
(d)(i)

Ms. Sharma decides to carry out a hypothesis test on the correlation coefficient to investigate whether increased weekly study hours are associated with higher exam scores. State why the hypothesis test should be one-tailed.

[1]
(d)(ii)

State the null and alternative hypotheses for this test.

[2]
(d)(iii)

The critical value for this test, at the 5% significance level, is 0.549. Ms. Sharma assumes that the population is bivariate normal. Determine whether there is significant evidence of a positive correlation between weekly study hours and exam scores. Justify your answer.

[2]
(e)(i)

Ms. Sharma wants to create a model to predict how changing weekly study hours might affect exam scores. To do this, she assumes that weekly study hours, XX, is the independent variable and the exam score, YY, is the dependent variable.

She first considers a linear model of the form Y=aX+bY = aX + b. Use Ms. Sharma's data to find the value of aa and of bb.

[1]
(e)(ii)

Interpret, referring to study hours and exam scores, what the value of aa represents.

[1]
(e)(iii)

Ms. Sharma then considers a quadratic model of the form Y=cX2+dX+eY = cX^2 + dX + e. Find the value of cc, of dd and of ee.

[1]
(e)(iv)

Find the coefficient of determination for each of the two models she considers.

[2]
(e)(v)

Hence compare the two models.

[1]
(e)(vi)

Ms. Sharma decides to use the coefficient of determination to choose between these two models. Comment on the validity of her decision.

[1]
(f)(i)

After presenting the results of her investigation, a colleague questions whether Ms. Sharma's sample is representative of all students in the school. A report states that the mean weekly study hours for all students in the school is 99 hours. Ms. Sharma decides to carry out a test to determine whether her sample could realistically be taken from a population with a mean of 99 hours. State the name of the test which Ms. Sharma should use.

[1]
(f)(ii)

State the null and alternative hypotheses for this test.

[1]
(f)(iii)

Perform the test, using a 5% significance level, and state your conclusion in context.

[3]

Question 15

MediumPaper 1 · calculator4 marks
(a)

The student body of Horizon University consists of 18 50018\,500 students.

A research team conducted a survey on student preferences for learning methods, selecting a random sample of 800800 students.

In this sample, 320320 students indicated a preference for online learning over traditional in-person classes.

Calculate an estimate for the total number of students at Horizon University who prefer online learning.

[2]
(b)

The research team decided to conduct a follow-up survey, but this time they opted for a stratified sample instead of a simple random sample.

[0]
(c)

Suggest two possible types of strata (apart from learning method preference) that would be sensible for the research team to use in their follow-up survey.

[2]

Question 16

HardPaper 3 · calculator28 marks
(a)(i)

(a) TechInnovate is considering collecting more data for their analysis.

(i) State one advantage of increasing the sample size.

[1]
(a)(ii)

(ii) State one disadvantage of increasing the sample size.

[1]
(b)

(b) The production manager at Plant Alpha recorded the time, in minutes, taken to produce a batch of electronic components for 10 randomly selected batches:

18.2,19.5,17.8,20.1,18.5,19.0,17.5,20.5,18.8,19.318.2, 19.5, 17.8, 20.1, 18.5, 19.0, 17.5, 20.5, 18.8, 19.3

Find the value of sn−1s_{n-1} for this sample from Plant Alpha.

[2]
(c)

(c) A manager claims that Plant Alpha's production times are more consistent than Plant Beta's. Given that the sample standard deviation (sn−1s_{n-1}) for Plant Beta's production times is 1.051.05 minutes, make one criticism of this claim.

[1]
(d)(i)

(d) TechInnovate wants to compare the mean production times of Plant Alpha and Plant Beta using a pooled t-test.

(i) State the condition regarding population variances required to use a pooled t-test.

[1]
(d)(ii)

(ii) Given that for Plant Beta, a sample of 1212 batches yielded a mean production time of xˉB=19.3\bar{x}_B = 19.3 minutes and a sample standard deviation of sB=1.05s_B = 1.05 minutes, state whether TechInnovate should use a pooled t-test in this case. Justify your answer.

[2]
(e)(i)

(e) TechInnovate believes Plant Alpha has a lower mean production time than Plant Beta.

(i) State appropriate null and alternative hypotheses for the pooled t-test.

[2]
(e)(ii)

(ii) Find the p-value.

[2]
(e)(iii)

(iii) Given that the test is carried out at the 5% significance level, state the appropriate conclusion in context. Justify your answer.

[2]
(f)(i)

(f) The company also investigates the relationship between operator experience (in years) and the number of defective items produced per day. A sample of 8 operators yielded the following data:

Operator Experience (years)Number of Defective Items
215
510
313
87
118
69
412
78

(i) Assuming all requirements are met, perform a test at the 5% significance level to determine if there is a linear correlation between operator experience and the number of defective items. State the hypotheses and justify your conclusion.

[4]
(f)(ii)

(ii) If the requirements for this test are not met, state an alternative test that could be used.

[1]
(g)

(g) For the data in (f.i), the equation of the least squares regression line of defective items (DD) on operator experience (EE) is D=−1.5E+18.25D = -1.5E + 18.25. Give, in context, an interpretation of the gradient −1.5-1.5 in this model.

[1]
(h)(i)

(h) TechInnovate uses a baseline model to predict the number of defective items (DpredD_{pred}) for a batch based on its size (SS): Dpred=0.5S+10D_{pred} = 0.5S + 10. The "Quality Deviation" (QQ) for a batch is defined as Q=Dpred−DactualQ = D_{pred} - D_{actual}. A positive Quality Deviation indicates better-than-expected quality.

(i) Show that for a batch of 150150 components from Plant Beta that produced 8080 defective items, the Quality Deviation is 5.05.0.

[2]
(h)(ii)

(ii) To compare quality control, samples of Quality Deviation scores were collected:

  • Plant Alpha: nQA=15n_{QA} = 15, xˉQA=4.5\bar{x}_{QA} = 4.5, sQA=1.2s_{QA} = 1.2
  • Plant Beta: nQB=18n_{QB} = 18, xˉQB=3.8\bar{x}_{QB} = 3.8, sQB=1.1s_{QB} = 1.1

Assuming that the appropriate requirements are met, use a pooled t-test at a 5% significance level to determine if the mean Quality Deviation is higher in Plant Alpha than in Plant Beta. Write down your null and alternative hypotheses and justify your conclusion.

[4]
(i)

(i) Using the results from parts (e) and (h.ii), explain how each plant could claim they are performing better than the other plant.

[2]

Question 17

MediumPaper 2 · calculator18 marks
(a)

A cognitive science researcher recorded the completion times, in seconds, for 30 participants solving a complex spatial puzzle. The data collected is as follows:

18202122222324252526272728282930303132323334353637383940425518 \quad 20 \quad 21 \quad 22 \quad 22 \quad 23 \quad 24 \quad 25 \quad 25 \quad 26 \quad 27 \quad 27 \quad 28 \quad 28 \quad 29 \quad 30 \quad 30 \quad 31 \quad 32 \quad 32 \quad 33 \quad 34 \quad 35 \quad 36 \quad 37 \quad 38 \quad 39 \quad 40 \quad 42 \quad 55

(a) Calculate the mean and standard deviation of the completion times. Comment on what these values indicate about the data.

[5]
(b)

(b) Determine the smallest value, the largest value, and the range of the completion times.

[3]
(c)

(c) Find the lower quartile (Q1Q_1), median, and upper quartile (Q3Q_3). Hence, calculate the interquartile range (IQR) and determine if there are any outliers in the data, justifying your answer.

[6]
(d)

(d) Draw a box-and-whisker plot for the data, clearly indicating any outliers.

[4]

Question 18

MediumPaper 1 · calculator10 marks
(a)

A farmer recorded the weights of 16 pumpkins harvested from his field. The five-number summary (excluding any outliers) and one outlier are given as follows:

Minimum weight: 4.54.5 kg

Lower Quartile (Q1Q_1): 6.06.0 kg

Median (Q2Q_2): 7.27.2 kg

Upper Quartile (Q3Q_3): 8.58.5 kg

Maximum weight (excluding outlier): 10.510.5 kg

One outlier was recorded at 13.013.0 kg.

(a) Write down the number of pumpkins that weighed more than 8.58.5 kg.

[1]
(b)

(b) Find the interquartile range (IQR) for the data.

[2]
(c)(i)

(c) Outliers are defined as values that fall outside the interval [Q1−1.5×IQR,Q3+1.5×IQR][Q_1 - 1.5 \times \text{IQR}, Q_3 + 1.5 \times \text{IQR}].

(i) Show that only one of these weights is an outlier.

[4]
(c)(ii)

(ii) Complete the box and whisker diagram below, including the outlier.

Placeholder for a box and whisker diagram with an x-axis ranging from 4 to 14, labelled 'Weight (kg)'. The box, whiskers, and outlier should be drawn according to the provided data.
[3]

Question 19

MediumPaper 1 · calculator6 marks
(a)

(a) State whether the data is discrete or continuous.

A quality control manager inspects items from a production line and records the number of defects found on each item. The data is presented in the following table:

Number of defects (xx)Number of items (ff)
015
125
235
3kk
410
55
[1]
(b)

(b) The mean number of defects per item is 22. Find the value of kk.

[4]
(c)

(c) The quality control manager divided the day's production into two shifts: morning and afternoon. She then randomly selected 1010 items from the morning shift's output and 1010 items from the afternoon shift's output for inspection. Identify the sampling technique used.

[1]

Question 20

MediumPaper 2 · calculator16 marks
(a)(i)

The scores of eight students in a national mathematics competition and the number of hours they spent studying are shown in the following table.

StudentScore (y)
A92
B88
C85
D80
E75
F72
G68
H65

(a)(i) For this data, find the upper quartile.

[2]
(a)(ii)

(a)(ii) For this data, find the interquartile range.

[2]
(b)

(b) Determine if Student A's score is an outlier for this data. Justify your answer.

[3]
(c)

A researcher is investigating the relationship between students' mathematics competition scores and their study hours to determine whether study hours can reasonably be used to predict a student's score.

The study hours of the students are shown in the table.

StudentStudy Hours (x)Score (y)
A4092
B3588
C2085
D3080
E1575
F2572
G1068
H1065

The researcher finds that, for this data, the Pearson's product moment correlation coefficient is r=0.35r = 0.35.

(c) State whether it would be appropriate for the researcher to use the equation of a regression line for yy on xx to predict a student's score. Justify your answer.

[2]
(d)(i)

The researcher then decides to find the Spearman's rank correlation coefficient for this data, and creates a table of ranks (lowest value = rank 1).

StudentScore RankStudy Hours Rank
A88
B77
C6a
D56
E43
F3b
G21.5
H1c

(d)(i) Write down the value of:

a,

[1]
(d)(ii)

(d)(ii) Write down the value of:

b,

[1]
(d)(iii)

(d)(iii) Write down the value of:

c.

[1]
(e)(i)

(e)(i) Find the value of the Spearman's rank correlation coefficient rsr_s.

[2]
(e)(ii)

(e)(ii) Interpret the value obtained for rsr_s.

[1]
(f)

(f) When calculating the ranks, the researcher incorrectly read Student A's score as 90. Explain why the value of the Spearman's rank correlation rsr_s does not change despite this error.

[1]

3 more Stats basics (population, 5 sampling techniques, outlier definition) questions in the app

Every answer is marked mark by mark, IB-style, and the AI tutor helps when you are stuck.

Where marks are lost

  • Rounding an intermediate value and then using it in a later part. Costs a mark every time, and AI's multi-part modelling questions give it more chances to happen than AA's shorter, more self-contained ones.
  • Using your own wrong value after failing a "show that." All follow through is withdrawn for the rest of that question.
  • Leaving an answer in calculator notation. Never accepted in a final answer, and AI's constant calculator use makes this the easiest slip in the whole subject.
Free. Every IB subject.
No card, no trial that runs out. Just a free account.
  • 50 marked answers a month
    Marked mark by mark, IB-style
  • Hints and mark schemes
    On every part of every question
  • 3,000+ questions
    All 6 subjects, SL and HL, mapped to the syllabus
  • Progress that adapts
    Your Study Profile picks what to practise next

Practise this topic as a session

Pick a difficulty and paper, and FourtyFive tracks your progress on this topic as you go.

or with email
FAQ

Questions,
answered.

Can't find what you're looking for? Email our student team.

What does Stats basics (population, 5 sampling techniques, outlier definition) cover in IB Maths AI?

Qualitative data is descriptive, non-numerical. Samples may not be representative of the entire population; minimise bias with random sampling. Stratified sampling divides the population into disjoint groups, then samples each group such that:.

Is Stats basics (population, 5 sampling techniques, outlier definition) SL or HL?

Both. SL and HL students study Stats basics (population, 5 sampling techniques, outlier definition) to the same depth.

How do I revise Stats basics (population, 5 sampling techniques, outlier definition) for IB Maths AI?

Start from the core idea: qualitative data is descriptive, non-numerical. In the exam: cheap marks and easy marks to lose. Naming a sampling technique is one mark, and saying why it is or is not appropriate here is another, and the second needs a reference to the actual context. Then practise exam-style questions, easiest first, writing out every step of your working before you check it.

How does FourtyFive help me practise Stats basics (population, 5 sampling techniques, outlier definition)?

FourtyFive has 23 Stats basics (population, 5 sampling techniques, outlier definition) questions. Every answer you write is marked mark by mark, IB-style, and you see where each mark was won or lost. Every part has a hint, the AI tutor helps you through the step you are stuck on, and your Study Profile picks what to practise next.

Is FourtyFive free for Stats basics (population, 5 sampling techniques, outlier definition) practice?

Yes. A free account gives you 50 marked answers a month, and you do not need a card to sign up.

Can I handwrite Stats basics (population, 5 sampling techniques, outlier definition) answers on an iPad?

Yes. In the FourtyFive iPad app you write your working by hand with Apple Pencil, the way you would on paper, and it is marked the same way.

Start with the IB question
bank built for you.

Free to start, no card needed. Thousands of syllabus-mapped questions, AI Examiner marking, your weakest topics first.