Skip to content
  1. IB Question Bank
  2. Maths AI
  3. Statistics & Probability
Topic 4.03 · SL and HL

Measures of central tendency / dispersion (std. Dev, var, IQR) + effect of constant changes to data: notes and practice questions

Summary
  • Measures of Central Tendency:
  • Mode: Most frequent value.
  • Median (Q2Q_2): Middle value of ordered data.
  • Mean (xˉ\bar{x} or μ\mu): Sum of data divided by number of data points.
  • Measures of Dispersion:
  • Quartiles (Q1,Q3Q_1, Q_3): Divide data into four equal sections.
  • Range: Maximum value - Minimum value.
  • Interquartile Range (IQR): Q3−Q1Q_3 - Q_1.
  • Variance (σ2\sigma^2): Mean of squared differences from the mean.
  • Standard Deviation (σ\sigma): Square root of variance; same units as data.
  • Mean Formulas:
  • Ungrouped data: xˉ=∑xin \bar{x} = \frac{\sum x_i}{n}
  • Frequency table: xˉ=∑fixinwheren=∑fi \bar{x} = \frac{\sum f_i x_i}{n} \quad \text{where} \quad n = \sum f_i
  • Mid-interval value (for grouped data): Upper boundary+Lower boundary2 \frac{\text{Upper boundary} + \text{Lower boundary}}{2}
  • Variance and Standard Deviation Formulas (GDC use expected):
  • Variance: σ2=∑fixi2n−μ2 \sigma^2 = \frac{\sum f_i x_i^2}{n} - \mu^2
  • Standard Deviation: σ=∑fixi2n−μ2 \sigma = \sqrt{\frac{\sum f_i x_i^2}{n} - \mu^2}
  • Comparing Data Sets:
  • With outliers: Use median (central tendency) and IQR (dispersion).
  • Symmetrical data: Use mean (central tendency) and standard deviation (dispersion).
  • **Linear Transformations (Adding/Subtracting a constant kk):**
  • Mean: New Mean=Old Mean+k \text{New Mean} = \text{Old Mean} + k
  • Standard Deviation, Variance, IQR, Range: Unchanged.
  • **Linear Transformations (Multiplying by a constant kk):**
  • Mean: New Mean=Old Mean×k \text{New Mean} = \text{Old Mean} \times k
  • Standard Deviation: New SD=Old SD×∣k∣ \text{New SD} = \text{Old SD} \times |k|
  • Variance: New Variance=Old Variance×k2 \text{New Variance} = \text{Old Variance} \times k^2
  • HL Formula Booklet (Expected Value & Variance):
  • E(aX+b)=aE(X)+b E(aX + b) = aE(X) + b
  • Var(aX+b)=a2Var(X) \text{Var}(aX + b) = a^2\text{Var}(X)
  • GDC Usage:
  • Input data into statistics mode for quartiles, SD, variance.
  • For grouped data, use mid-interval values as xx-values.
  • Always check GDC output for logical consistency.
  • Exam Pitfall: Adding/subtracting a constant does not change measures of dispersion (SD, variance, IQR, range).
  • Grouped Data Estimates: Indicate estimates (e.g., by rounding to 3 s.f.) as values are not exact.

How it is examined

The quartile warning is worth taking seriously when marking: a hand-calculated quartile can legitimately differ from the GDC's, so a mark scheme should accept both unless the question forces one method. The constant-change results are a recurring two-mark question and are pure reasoning, no calculation. Estimating a mean from grouped data needs mid-interval values, and using the lower bound instead is the standard error.

Given in the booklet

The mean of a set of data, xˉ=∑i=1kfixin\bar{x} = \dfrac{\sum_{i=1}^{k} f_i x_i}{n}, where n=∑i=1kfin = \sum_{i=1}^{k} f_i.

Key ideas
  • Measures of central tendency: mean, median and mode.
  • Estimation of the mean from grouped data.
  • The modal class.
  • Measures of dispersion: interquartile range, standard deviation and variance.

Linking questions

  • Other contexts: comparing variation and spread in populations, human or natural, for example agricultural crop data, social indicators, reliability and maintenance.
  • Links to other subjects: descriptive statistics (sciences, individuals and societies); the consumer price index (economics).
  • International-mindedness: the benefits of sharing and analysing data from different countries; discussion of the different formulae for variance.
  • TOK: could mathematics make alternative, equally true, formulae? What does that tell us about mathematical truths? Does the use of statistics lead to an over-emphasis on attributes that can be measured easily over those that cannot?

Practice questions

55 questions · 1 easy · 41 medium · 13 hard
Showing 20 of 20

Question 1

EasyPaper 1 · calculator3 marks

A tech company conducts tests on the battery life (in hours) of its new smartphone model. The results show that the battery life has a mean of 2525 hours and a standard deviation of 2.52.5 hours.

Following a software update, every smartphone receives an additional 3.53.5 hours of battery life. Write down the new mean battery life and the new standard deviation of the battery life.

Question 2

MediumPaper 1 · calculator4 marks
(a)

A group of ten students recorded the number of hours they spent studying for their final IB Mathematics exam in the week leading up to it.

Their results are shown below:

5, 7, 6, 8, 5, 9, 7, 6, 5, 10

For this data set, find the value of

(a) the mode.

[1]
(b)

A group of ten students recorded the number of hours they spent studying for their final IB Mathematics exam in the week leading up to it.

Their results are shown below:

5, 7, 6, 8, 5, 9, 7, 6, 5, 10

For this data set, find the value of

(b) the mean.

[2]
(c)

A group of ten students recorded the number of hours they spent studying for their final IB Mathematics exam in the week leading up to it.

Their results are shown below:

5, 7, 6, 8, 5, 9, 7, 6, 5, 10

For this data set, find the value of

(c) the standard deviation.

[1]

Question 3

HardPaper 1 · calculator11 marks
(a)

(a) A botanist is studying a rare species of flower. They record the number of petals (pp) on a sample of these flowers. The data is presented in the frequency table below:

pp

5

6

7

8

9

10

Frequency

3

8

20

15

7

2

Find an unbiased estimate of the population mean number of petals for this rare flower species.

[2]
(b)

(b) Find an unbiased estimate of the population variance of the number of petals for this rare flower species.

[3]
(c)(i)

(c) A botanist suspects that the average number of petals for this rare species is different from the average of 7.0 petals observed in a more common, related species. She sets up a hypothesis test with the null hypothesis H0:μ=7.0H_0: \mu = 7.0.

(i) State the alternative hypothesis.

[1]
(c)(ii)

(ii) Given that all assumptions for this test are satisfied, carry out an appropriate hypothesis test. State and justify your conclusion, using a 10% significance level.

[5]

Question 4

MediumPaper 1 · calculator9 marks
(a)

A software development company tracked the completion times of 150 projects. The cumulative frequency graph shows the completion times obtained by the projects.

A cumulative frequency graph. The x-axis is Project completion time (days) from 0 to 100. The y-axis is Cumulative number of projects from 0 to 150. The curve starts at (0,0) and increases to (100,150), showing an S-shape.

Find the median completion time of the projects.

[1]
(b)

The projects were assigned a performance tier from 1 to 5, depending on the completion time. The number of projects receiving each tier is shown in the following table.

Tier12345
Number of projects81530pq

Find an expression for pp in terms of qq.

[2]
(c)(i)

The mean performance tier for these projects is 3.5.

Find the number of projects that obtained a tier 5.

[4]
(c)(ii)

Find the minimum completion time needed to obtain a tier 5.

[2]

Question 5

HardPaper 2 · calculator12 marks
(a)

The following tables show the mean monthly rainfall, by month, in two different regions: Green Valley and Sunstone Coast.

Green Valley

Month | Mean monthly rainfall (mm)

---|---

January | 8080

February | 7575

March | 9090

April | 6060

May | 5050

June | 3030

July | 2020

August | 2525

September | 4040

October | 7070

November | 9595

December | 8585

Sunstone Coast

Month | Mean monthly rainfall (mm)

---|---

January | 4040

February | 3535

March | 4545

April | 5050

May | 6060

June | 7070

July | 8080

August | 7575

September | 6565

October | 5555

November | 4848

December | 4242

(a) Find the mean monthly rainfall over the course of the year for Green Valley.

[2]
(b)

(b) Find the standard deviation of the monthly rainfall in Green Valley.

[2]
(c)

(c) Find the mean monthly rainfall over the course of the year for Sunstone Coast.

[2]
(d)

(d) Find the standard deviation of the monthly rainfall in Sunstone Coast.

[2]
(e)

(e) By referring directly to your answers from parts (a) -(d), make contextual comparisons about the monthly rainfall in Green Valley and Sunstone Coast throughout the year.

[4]

Question 6

MediumPaper 1 · calculator5 marks
(a)

A factory conducts quality control checks on the diameter (in mm) of components produced by Production Line A. A sample of measurements is recorded as:

18.2, 19.5, 20.1, 20.3, 20.5, 20.6, 20.8, 21.0, 21.2, 21.5, 22.0, 22.1, 22.3, 22.5, 23.0

For these data, the lower quartile is 20.3 mm and the upper quartile is 22.1 mm.

Show that a component with a diameter of 18.2 mm would not be considered an outlier.

[3]
(b)

Another production line, Line B, also produces similar components. The box and whisker diagram below displays the diameters (in mm) of a sample of components from Line B.

Box and whisker plot for Production Line B diameters. The plot shows: Minimum = 17.0, Lower Quartile (Q1) = 20.0, Median (Q2) = 21.5, Upper Quartile (Q3) = 23.0, Maximum = 24.5.

A quality control manager reviews the box and whisker diagrams for both lines and suggests that Production Line B generally produces components with larger diameters.

With reference to the box and whisker diagrams for Line A (from part (a) ) and Line B, state one aspect that may support the manager's opinion and one aspect that may counter it.

[2]

Question 7

HardPaper 2 · calculator22 marks
(a)

A logistics company recorded the delivery times (in minutes) for a large batch of packages. The data is grouped in the frequency table below:

Delivery Time (minutes) | Frequency

---|---

20≤t<2520 \le t < 25 | 88

25≤t<3025 \le t < 30 | 2222

30≤t<3530 \le t < 35 | 5555

35≤t<4035 \le t < 40 | 9898

40≤t<4540 \le t < 45 | 135135

45≤t<5045 \le t < 50 | 110110

50≤t<5550 \le t < 55 | 7070

55≤t<6055 \le t < 60 | 3535

60≤t<6560 \le t < 65 | 1515

65≤t<7065 \le t < 70 | 22

(a) Calculate estimates of the mean and standard deviation of the delivery times.

[4]
(b)

(b) Construct a cumulative frequency table for the data, and use it to draw a cumulative frequency curve.

Cumulative frequency curve placeholder. X-axis: Delivery Time (minutes), Y-axis: Cumulative Frequency.
[4]
(c)(i)

(c) Use your graph to estimate:

(i) the median delivery time

[2]
(c)(ii)

(ii) the lower and upper quartile of the delivery times

[2]
(c)(iii)

(iii) the interquartile range

[1]
(c)(iv)

(iv) the 8585th percentile of delivery times.

[2]
(d)

(d) Draw a box-and-whisker plot of the data.

Box and whisker plot placeholder. X-axis: Delivery Time (minutes).
[3]
(e)

(e) Determine, with reasons, whether any customers could be considered outliers.

[4]

Question 8

MediumPaper 1 · calculator8 marks
(a)

A group of 160 participants completed a fitness challenge. The cumulative frequency graph shows the points obtained by the participants.

Cumulative frequency graph of fitness challenge points

Find the median of the points obtained.

[1]
(b)

The participants were awarded a performance grade from 1 to 5, depending on the points obtained in the challenge.

The number of participants receiving each grade is shown in the following table.

Grade12345
Number of participants81530ab

Find an expression for aa in terms of bb.

[2]
(c)(i)

The mean grade for these participants is 3.5.

Find the number of participants who obtained a grade 5.

[3]
(c)(ii)

Find the minimum points needed to obtain a grade 5.

[2]

Question 9

HardPaper 2 · calculator12 marks
(a)

A survey was conducted to investigate the daily screen time (in minutes) of 80 students. The results are presented in the frequency table below.

Time (tt minutes) | Frequency

---|---

0≤t<300 \le t < 30 | 66

30≤t<6030 \le t < 60 | 1010

60≤t<9060 \le t < 90 | 2222

90≤t<12090 \le t < 120 | 2020

120≤t<150120 \le t < 150 | 1212

150≤t<180150 \le t < 180 | 77

180≤t<210180 \le t < 210 | 33

(a) State the modal class.

[1]
(b)

(b) Find the class in which the median time lies.

[2]
(c)

(c) Construct a cumulative frequency table for this data.

[3]
(d)

(d) Sketch the cumulative frequency curve.

[2]
(e)

(e) Use your curve to find estimates for the median and interquartile range.

[4]

Question 10

MediumPaper 1 · calculator19 marks
(a)(i)

[Maximum mark: 19]

A tech company recorded the time (in minutes) 180 customers spent completing a new online feedback survey. The data was compiled into the following cumulative frequency graph.

Cumulative frequency graph showing time in minutes on x-axis and cumulative frequency on y-axis (from 0 to 180)

(a) Use the graph to find

(i) the median time;

[4]
(a)(ii)

(ii) the lower quartile;

[4]
(a)(iii)

(iii) the upper quartile;

[4]
(a)(iv)

(iv) the interquartile range.

[4]
(b)

Sarah completed the survey in 1.5 minutes.

(b) Determine whether Sarah's time is an outlier.

[3]

Question 11

HardPaper 2 · calculator21 marks
(a)

Dr. Anya Sharma, a sports scientist, is investigating the relationship between training habits and performance in junior athletes. She wants to collect data on the weekly training hours of junior swimmers. She decides to interview every 5th swimmer entering the training facility until she has a sample of 50 swimmers.

State the sampling method Dr. Sharma has used.

[1]
(b)

Dr. Sharma constructed the following box and whisker diagram to show the weekly training hours (in hours) of a sample of junior swimmers.

A box and whisker diagram showing weekly training hours. The minimum is 2, the first quartile (Q1) is 4, the median is 6, the third quartile (Q3) is 9, and the maximum is 12.

Write down the median weekly training hours.

[1]
(c)

Calculate the interquartile range for the weekly training hours.

[2]
(d)

One swimmer in the sample reported training for 15 hours per week. Dr. Sharma believes this swimmer's training time is not an outlier.

Determine whether Dr. Sharma is correct. Support your reasoning.

[4]
(e)

Dr. Sharma also collected data on the average weekly training hours (xx) and the competition score (yy) for a group of athletes. These data are represented on the scatter diagram.

A scatter diagram showing competition score (y-axis from 0 to 120) versus weekly training hours (x-axis from 0 to 25). The points show a general negative correlation, with data points roughly between 5 and 20 hours.

Describe the correlation between weekly training hours and competition score.

[1]
(f)

Dr. Sharma correctly calculates the equation of the regression line yy on xx for these athletes to be y=−2.5x+110y = -2.5x + 110. She uses the equation to estimate the competition score for an athlete who trains 3 hours per week.

Find the competition score calculated by Dr. Sharma.

[2]
(g)

State whether it is valid to use the regression line yy on xx for Dr. Sharma's estimate in part (f). Give a reason for your answer, assuming the original data for training hours ranged from 5 to 20 hours.

[2]
(h)

Dr. Sharma investigated the relationship between an athlete's national competition rank and their average daily protein intake (in grams). She collected data for eight athletes, as shown in the table.

AthleteABCDEFGH
Competition Rank (RcompR_{comp})12345678
Protein Intake (g) (PintakeP_{intake})180150200160140190170130

Dr. Sharma intends to analyse the data using Spearman's rank correlation coefficient, rsr_s.

Copy and complete the information in the following table.

AthleteABCDEFGH
Rank - Competition Rank1
Rank - Protein Intake
[2]
(i)(i)

Calculate the value of rsr_s.

[3]
(i)(ii)

Interpret your result.

[3]

Question 12

MediumPaper 1 · calculator4 marks
(a)

[Maximum mark: 4]

A salesperson records the number of items sold per day over several weeks. The data is presented in the following frequency distribution table.

Number of items sold (xx)Frequency (ff)
13
25
38
4p
56
64

| Frequency (f)

------------------------|---------------

1 | 3

2 | 5

3 | 8

4 | p

5 | 6

6 | 4

)

For this distribution, the mean number of items sold per day is 3.8.

(a) Write down the total number of days the salesperson recorded data in terms of p.

[1]
(b)

(b) Calculate the value of p.

[3]

Question 13

HardPaper 2 · calculator15 marks
(a)

A quality control manager at a manufacturing plant wants to assess the consistency of a new batch of electronic components. He decides to test ten components, ensuring that five are selected from Production Line A and five from Production Line B. The manager instructs the supervisors of each line to provide the required number of components from their current production.

(a) Name the type of sampling that best describes the method used by the quality control manager.

[1]
(b)(i)

The weights, in grams, of the ten components selected for the test are:

148,153,161,155,142,160,150,157,145,159148, 153, 161, 155, 142, 160, 150, 157, 145, 159.

(b) For these ten components, find

(i) the mean weight.

[2]
(b)(ii)

(ii) the standard deviation of the weights.

[2]
(c)

The target weight for these components is 155155 g. The manager is concerned that the components might be consistently underweight. Perform an appropriate test at the 10%10\% significance level to see if the mean weight of the components produced is less than the target weight. It can be assumed that the weights come from a normal population.

[5]
(d)

State one reason why the test performed in part (c) might not be valid.

[1]
(e)(i)

Two additional components are tested at a later date. The mean weight for all twelve components is 154.5154.5 g and the standard deviation is 7.27.2 g.

For further analysis, a 'quality score' for the twelve components is obtained by multiplying the weights by 1.51.5 and subtracting 5050.

(e) For the twelve components, find

(i) their mean quality score.

[2]
(e)(ii)

(ii) the standard deviation of their quality score.

[2]

Question 14

MediumPaper 1 · calculator7 marks
(a)

A tech company launched two new smartphone models, "Voyager" and "Explorer". They collected customer satisfaction scores (out of 100) from a large sample of users for both models. The results are summarized in the following box and whisker diagram.

Box and whisker diagram comparing customer satisfaction scores for Model Voyager and Model Explorer. X-axis from 50 to 100. Model Voyager: min 50, Q1 60, median 70, Q3 90, max 100. Model Explorer: min 55, Q1 70, median 80, Q3 90, max 100.

Identify which two of the following statements must be true according to the box and whisker diagram. Indicate your choices by placing tick marks in the second column of the following table.

Statement | True (✓)

---|---

The satisfaction scores for Model Voyager are normally distributed. |

A higher percentage of customers gave a score less than 70 for Model Voyager than for Model Explorer. |

A higher percentage of customers gave a score greater than 90 for Model Explorer than for Model Voyager. |

The interquartile range for Model Explorer is less than the interquartile range for Model Voyager. |

[2]
(b)

A product manager believes there is no significant difference in the average customer satisfaction scores between the two models. She plans to conduct a t-test at the 10% significance level. Write down the null and alternative hypotheses for her test.

[2]
(c)

The t-test yielded a p-value of 0.0783. Find the p-value for her test.

[1]
(d)

Write down the conclusion to the test. Give a reason for your answer.

[2]

Question 15

HardPaper 2 · calculator21 marks
(a)(i)

The lifespans, tt, of 250 LED light bulbs are recorded in the following table.

Lifespan (hours)Frequency
0≤t<10000 \le t < 100020
1000≤t<15001000 \le t < 150060
1500≤t<20001500 \le t < 200090
2000≤t<25002000 \le t < 250055
2500≤t<30002500 \le t < 300025

This table is used to create a cumulative frequency graph.

Write down the mid-interval value of the class 0≤t<10000 \le t < 1000.

[1]
(a)(ii)

Calculate an estimate of the mean lifespan of the 250 light bulbs.

[3]
(b)

Use the cumulative frequency curve (which would be provided in an exam) to estimate the interquartile range. Assume the lower quartile (Q1Q_1) is 13001300 hours and the upper quartile (Q3Q_3) is 21502150 hours.

[3]
(c)

A light bulb from the data set had a lifespan of 34003400 hours.

Use your answer to part (b) to estimate whether this light bulb's lifespan is an outlier for this data. Justify your answer.

[3]
(d)

It is believed that the lifespans of these LED light bulbs follow a normal distribution with mean 17401740 hours and standard deviation 450450 hours.

It is decided to perform a χ2\chi^2 goodness of fit test on the data to determine whether this sample of 250 light bulbs could have plausibly been drawn from an underlying distribution N(1740,4502)N(1740, 450^2).

Write down the null and the alternative hypotheses for the test.

[2]
(e)(i)

As part of the test, the following table is created.

Lifespan of light bulb (hours)Observed frequencyExpected frequency
t<1000t < 10002014.0
1000≤t<15001000 \le t < 15006060.1
1500≤t<20001500 \le t < 200090a
2000≤t<25002000 \le t < 25005560.1
t≥2500t \ge 250025b

Find the value of aa and the value of bb. Give your answers to one decimal place.

[5]
(e)(ii)

Hence, perform the test to a 5% significance level, clearly stating the conclusion in context.

[4]

Question 16

MediumPaper 1 · calculator7 marks
(a)(i)

A logistics company recorded the weights, in kilograms (kg), of six parcels dispatched in a single hour:

1.25 kg, 1.80 kg, 1.50 kg, 1.25 kg, 2.10 kg, 1.60 kg

For these six parcels, find

(i) the mean weight.

[2]
(a)(ii)

(ii) the median weight.

[1]
(a)(iii)

(iii) the modal weight.

[1]
(a)(iv)

(iv) the range of the weights.

[2]
(b)

A new parcel is added to the shipment. Its weight is measured as 1.75 kg to the nearest 10 grams.

Write down the shortest possible weight of this new parcel.

[1]

Question 17

HardPaper 2 · calculator21 marks
(a)

A company manufactures specialized medical devices. Each device consists of a main circuit board and a protective casing. The weight of the circuit board, CC, is normally distributed with a mean of 120120 g and a standard deviation of 55 g. The weight of the protective casing, PP, is normally distributed with a mean of 3030 g and a standard deviation of 22 g. The weights of the circuit board and the casing are independent.

Find the probability that a randomly chosen complete device has a total weight of less than 145145 g.

[5]
(b)

A batch of 1010 such devices is to be packed into a container. The container has a maximum weight capacity of 15101510 g. The weight of each device is independent.

Find the probability that the total weight of the 1010 devices is greater than the capacity of the container.

[4]
(c)(i)

The company sources a critical microchip component from two different suppliers, Supplier X and Supplier Y. An engineer claims that microchips from Supplier Y have a lower response time than those from Supplier X. To test this claim, a random sample is taken from each supplier.

The eight microchips in the sample from Supplier X have response times, in milliseconds (ms), of:

19.65,19.65,19.79,20.75,20.97,21.15,22.28,22.3719.65, 19.65, 19.79, 20.75, 20.97, 21.15, 22.28, 22.37.

Find the

(i) mean response time for the sample from Supplier X.

[3]
(c)(ii)

Find the

(ii) unbiased estimate of the population variance for the sample from Supplier X.

[3]
(d)

The seven microchips in the sample from Supplier Y have a mean response time of 19.519.5 ms and an unbiased estimate of the population standard deviation (sn−1s_{n-1}) of 1.051.05 ms.

Perform a suitable test, at the 5%5\% significance level, to test the engineer's claim that microchips from Supplier Y have a lower response time than those from Supplier X. You may assume the response times of microchips from each supplier are normally distributed with equal population variance.

[6]

Question 18

MediumPaper 1 · calculator6 marks
(a)(i)

(a) The formula for converting daily steps, SS, to a fitness score, FF, is given by F=0.05S+10F = 0.05S + 10.

(i) Find a formula for converting a fitness score, FF, back to daily steps, SS.

[2]
(a)(ii)

(ii) A user achieved a fitness score of 75. Calculate the number of daily steps they took.

[1]
(b)(i)

(b) Over a month, the mean daily steps recorded by a group of users was 8500 steps with a standard deviation of 1200 steps.

For the same group, find

(i) the mean daily fitness score.

[1]
(b)(ii)

(ii) the standard deviation of the daily fitness scores.

[2]

Question 19

HardPaper 2 · calculator21 marks
(a)(i)

(a) The scores, ss, of 200 students on a mathematics test are recorded in the following table.

Score (ss)Frequency
20≤s<4020 \le s < 4015
40≤s<6040 \le s < 6035
60≤s<8060 \le s < 8060
80≤s<10080 \le s < 10050
100≤s<120100 \le s < 12030
120≤s<140120 \le s < 14010

(i) Write down the mid-interval value of 60≤s<8060 \le s < 80.

[3]
(a)(ii)

(ii) Calculate an estimate of the mean score of the 200 students.

[3]
(b)

(b) The data from this table is used to create a cumulative frequency graph. From this graph, the first quartile (Q1Q_1) is estimated to be 6060 and the third quartile (Q3Q_3) is estimated to be 9696.

Use these values to estimate the interquartile range (IQR).

[2]
(c)

(c) A student, Elara, scored 155155 on the test.

Use your answer to part (b) to estimate whether Elara's score is an outlier for this data. Justify your answer.

[3]
(d)

(d) It is believed that the scores of students on this mathematics test follow a normal distribution with mean 77.577.5 and standard deviation 2020.

It is decided to perform a χ2\chi^2 goodness of fit test on the data to determine whether this sample of 200 students could have plausibly been drawn from an underlying distribution N(77.5,202)N(77.5, 20^2).

Write down the null and the alternative hypotheses for the test.

[2]
(e)

(e) As part of the test, the following table is created, where some categories have been combined to ensure expected frequencies are not too low.

Score (ss)Observed FrequencyExpected Frequency
s<40s < 40156.08
40≤s<6040 \le s < 603532.08
60≤s<8060 \le s < 8060a
80≤s<10080 \le s < 1005063.99
s≥100s \ge 10040b

(i) Find the value of aa and the value of bb.

(ii) Hence, perform the test to a 5% significance level, clearly stating the conclusion in context.

[8]

Question 20

MediumPaper 1 · calculator9 marks
(a)(i)

(a) A quality control inspector at a beverage company takes a random sample of eight bottles of orange juice from a production line. The measured volumes, in ml, are:

298.5, 301.2, 299.1, 300.8, 298.9, 301.5, 299.7, 300.3

(i) Find an unbiased estimate for the mean volume of orange juice in a bottle from this production line.

[2]
(a)(ii)

(ii) Calculate a 95% confidence interval for the population mean volume. Give your answer to four significant figures.

[4]
(b)

(b) State one assumption you have made in order for your interval to be valid.

[1]
(c)

(c) The label on each bottle states: "Volume: 300 ml".

Using your answer to part (a)(ii), briefly comment on the claim on the label.

[2]

35 more Measures of central tendency / dispersion (std. Dev, var, IQR) + effect of constant changes to data questions in the app

Every answer is marked mark by mark, IB-style, and the AI tutor helps when you are stuck.

Where marks are lost

  • Using your own wrong value after failing a "show that." All follow through is withdrawn for the rest of that question.
  • Leaving an answer in calculator notation. Never accepted in a final answer, and AI's constant calculator use makes this the easiest slip in the whole subject.
Free. Every IB subject.
No card, no trial that runs out. Just a free account.
  • 50 marked answers a month
    Marked mark by mark, IB-style
  • Hints and mark schemes
    On every part of every question
  • 3,000+ questions
    All 6 subjects, SL and HL, mapped to the syllabus
  • Progress that adapts
    Your Study Profile picks what to practise next

Practise this topic as a session

Pick a difficulty and paper, and FourtyFive tracks your progress on this topic as you go.

or with email
FAQ

Questions,
answered.

Can't find what you're looking for? Email our student team.

What does Measures of central tendency / dispersion (std. Dev, var, IQR) + effect of constant changes to data cover in IB Maths AI?

Measures of Central Tendency:. Mode: Most frequent value. Median (Q_2): Middle value of ordered data.

Is Measures of central tendency / dispersion (std. Dev, var, IQR) + effect of constant changes to data SL or HL?

Both. SL and HL students study Measures of central tendency / dispersion (std. Dev, var, IQR) + effect of constant changes to data to the same depth.

How do I revise Measures of central tendency / dispersion (std. Dev, var, IQR) + effect of constant changes to data for IB Maths AI?

Start from the core idea: measures of Central Tendency:. In the exam: the quartile warning is worth taking seriously when marking: a hand-calculated quartile can legitimately differ from the GDC's, so a mark scheme should accept both unless the question forces one method. The constant-change results are a recurring two-mark question and are pure reasoning, no calculation. Then practise exam-style questions, easiest first, writing out every step of your working before you check it.

How does FourtyFive help me practise Measures of central tendency / dispersion (std. Dev, var, IQR) + effect of constant changes to data?

FourtyFive has 55 Measures of central tendency / dispersion (std. Dev, var, IQR) + effect of constant changes to data questions. Every answer you write is marked mark by mark, IB-style, and you see where each mark was won or lost. Every part has a hint, the AI tutor helps you through the step you are stuck on, and your Study Profile picks what to practise next.

Is FourtyFive free for Measures of central tendency / dispersion (std. Dev, var, IQR) + effect of constant changes to data practice?

Yes. A free account gives you 50 marked answers a month, and you do not need a card to sign up.

Can I handwrite Measures of central tendency / dispersion (std. Dev, var, IQR) + effect of constant changes to data answers on an iPad?

Yes. In the FourtyFive iPad app you write your working by hand with Apple Pencil, the way you would on paper, and it is marked the same way.

Start with the IB question
bank built for you.

Free to start, no card needed. Thousands of syllabus-mapped questions, AI Examiner marking, your weakest topics first.