Skip to content
  1. IB Question Bank
  2. Maths AI
  3. Statistics & Probability
Topic 4.04 · SL and HL

Linear Correlation of bivariate data (scatter diagrams, lines of best fit, Pearson) + regression line interpretation: notes and practice questions

Summary
  • Bivariate Data: Data on two variables, paired to examine relationships.
  • Correlation vs. Causation: Correlation does not imply causation.
  • Linear Correlation Types (Scatter Diagram):
  • Strong Positive: Tightly clustered, upward line.
  • Weak Positive: Upward trend, loosely scattered.
  • Strong Negative: Tightly clustered, downward line.
  • Weak Negative: Downward trend, loosely scattered.
  • No Correlation: Scattered, no linear pattern.
  • **Pearson's Product-Moment Correlation Coefficient (PMCC), rr:* Measures strength and direction of linear* relationship.
  • Range: −1≤r≤1-1 \leq r \leq 1.
  • r=1r=1: Perfect positive linear correlation.
  • r=−1r=-1: Perfect negative linear correlation.
  • r=0r=0: No linear correlation.
  • Closer to 11 or −1-1 indicates stronger linear correlation.
  • Pearson's PMCC Formula:

r=SxySxSy r = \frac{S_{xy}}{S_x S_y}
Where:
Sxy=∑i=1nxiyi−1n(∑i=1nxi)(∑i=1nyi) S_{xy} = \sum_{i=1}^{n} x_i y_i - \frac{1}{n} \left( \sum_{i=1}^{n} x_i \right) \left( \sum_{i=1}^{n} y_i \right)
Sx=∑i=1nxi2−1n(∑i=1nxi)2 S_x = \sqrt{ \sum_{i=1}^{n} x_i^2 - \frac{1}{n} \left( \sum_{i=1}^{n} x_i \right)^2 }
Sy=∑i=1nyi2−1n(∑i=1nyi)2 S_y = \sqrt{ \sum_{i=1}^{n} y_i^2 - \frac{1}{n} \left( \sum_{i=1}^{n} y_i \right)^2 }

  • Regression Line Equation (y on x):

y=ax+b y = ax + b

  • **Gradient (aa):** Change in yy for each unit change in xx.
  • Positive aa: yy increases by aa for unit xx increase.
  • Negative aa: yy decreases by ∣a∣|a| for unit xx increase.
  • **Y-intercept (bb):** Value of yy when x=0x=0.
  • Mean Point: (xˉ,yˉ)(\bar{x}, \bar{y}) always lies on the regression line.
  • Drawing Line of Best Fit by Eye: Plot data, calculate and plot (xˉ,yˉ)(\bar{x}, \bar{y}), draw line through mean point following trend.
  • Spearman's Rank Correlation Coefficient: Calculate PMCC on ranked data (tied values get mean rank).
  • GDC Use: Enter data into lists, use linear regression (`ax + b`) to find a,b,ra, b, r. Store full a,ba, b values for predictions to avoid rounding errors.
  • HL Extension: Non-linear regression (quadratic, cubic, exponential y=abxy = ab^x, power y=axby = ax^b, sinusoidal). Linearise exponential (ln⁡y\ln y vs xx) and power (ln⁡y\ln y vs ln⁡x\ln x) relationships using logarithms.
  • Outliers: Bivariate outliers are distinct from univariate outliers.
  • Rounding Errors: Use exact aa and bb values for intermediate prediction steps.
  • Contextual Logic: Interpretations and predictions must be reasonable for the given situation.

How it is examined

One of the highest-frequency subtopics in the course. The pattern is: calculate rr, describe the correlation in words, find the regression equation, use it to predict, then comment on the reliability of that prediction. The last two marks are where students lose out, because "the prediction is unreliable" needs a reason, usually extrapolation beyond the data range or a weak rr. Interpreting aa and bb in context (a rate and a starting value) is a separate mark and students give generic answers. Predicting xx from yy using the yy on xx line is the specific error the guide calls out.

Given in the booklet

The formula for Pearson's rr and for the regression line of yy on xx.

Key ideas
  • Linear correlation of bivariate data.
  • Pearson's product-moment correlation coefficient, rr.
  • Scatter diagrams, and lines of best fit by eye passing through the mean point.
  • The equation of the regression line of yy on xx.

Linking questions

  • Other contexts: linear regressions where correlation exists between two variables. Exploring cause and dependence for categorical variables, for example what factors political persuasion might depend on.
  • Links to other subjects: curves of best fit, correlation and causation (sciences); scatter graphs (geography).
  • Aim 8: the correlation between smoking and lung cancer was discovered using mathematics, and science then had to justify the cause.
  • TOK: correlation and causation. Can we have knowledge of cause and effect relationships given that we can only observe correlation? What factors affect the reliability and validity of mathematical models in describing real-life phenomena?

Practice questions

36 questions · 2 easy · 27 medium · 7 hard
Showing 20 of 20

Question 1

EasyPaper 1 · calculator5 marks
(a)

A marine biologist is analysing data collected from various ocean environments. For each scenario described below, decide which of the following statements best represents the relationship between the two variables shown in the scatter diagram.

Statements:

I. Strong positive linear correlation

II. Weak positive linear correlation

III. No correlation

IV. Weak negative linear correlation

V. Strong negative linear correlation

(a) The scatter diagram represents the relationship between water temperature and the metabolic rate of a certain fish species.

Scatter diagram showing data points that are tightly clustered around a line that slopes upwards from left to right
[1]
(b)

(b) The scatter diagram represents the relationship between ocean depth and the amount of light penetration.

Scatter diagram showing data points that generally trend downwards from left to right, but are quite spread out
[1]
(c)

(c) The scatter diagram represents the relationship between the salinity of the water and the number of barnacles on a specific rock.

Scatter diagram showing data points that are randomly scattered across the graph with no apparent pattern or trend
[1]
(d)

(d) The scatter diagram represents the relationship between the level of ocean pollution and the health index of a coral reef.

Scatter diagram showing data points that are very closely clustered around a line that slopes downwards from left to right
[1]
(e)

(e) The scatter diagram represents the relationship between the density of plankton and the population size of a specific small fish species.

Scatter diagram showing data points that generally trend upwards from left to right, but are noticeably spread out
[1]

Question 2

MediumPaper 1 · calculator6 marks
(a)

Clara, a small business owner, believes there is a linear relationship between the daily advertising spend and the weekly sales for her new product.

To investigate this, she recorded the daily advertising spend, x hundreds of dollars, and the weekly sales, y thousands of dollars, for eight consecutive weeks. Her results are presented in the following table and scatter diagram.

x, hundreds of dollarsy, thousands of dollars
511.7
812.6
1015.0
1217.5
1516.6
1818.4
2022.4
2222.4
A scatter diagram showing weekly sales (y) on the y-axis and daily advertising spend (x) on the x-axis, with 8 data points plotted. The y-axis ranges from 10 to 25, and the x-axis from 0 to 25. The points generally show an increasing trend.

Clara consulted a business analytics textbook and found the following guidelines for interpreting the Pearson's product-moment correlation coefficient, r:

Value of ∣r∣\lvert r \rvertDescription of the correlation
0≤∣r∣<0.40 \le \lvert r \rvert < 0.4weak
0.4≤∣r∣<0.80.4 \le \lvert r \rvert < 0.8moderate
0.8≤∣r∣≤10.8 \le \lvert r \rvert \le 1strong

For this data, find the value of the Pearson's product-moment correlation coefficient, r.

[2]
(b)

Comment on your answer to part (a), using the information that Clara found.

[1]
(c)

Write down the equation of the regression line of y on x, in the form y=ax+by = ax + b.

[1]
(d)

Clara is considering a daily advertising spend of 17 hundreds of dollars for the next week.

Use the equation of the regression line to estimate the weekly sales for this advertising spend.

[2]

Question 3

HardPaper 2 · calculator16 marks
(a)

A horticulturalist is studying the relationship between the average daily temperature (xx, in ∘C^\circ\text{C}) and the weekly growth (yy, in mm per week) of a new plant species. The results from a sample of plants are shown in the table below.

Temperature (xx, ∘C^\circ\text{C})181820202222242426262828303032323434363638384040
Growth (yy, mm/week)252529293434373740404545484852525555595963636767

Find Pearson's product moment correlation coefficient, rr, for this data.

[2]
(b)

Describe the correlation between the average daily temperature and the weekly plant growth.

[2]
(c)

Explain why it is appropriate to find the regression line of yy on xx.

[2]
(d)

Find the regression line of yy on xx. State the domain on which it has been defined.

[4]
(e)

Flora's plant was kept at an average daily temperature of 27 ∘C27~^\circ\text{C} but its growth was not recorded. The horticulturalist uses the regression line to estimate the weekly growth Flora's plant would have obtained.

Find the estimated weekly growth for Flora's plant. Give your answer correct to 22 significant figures.

[3]
(f)

During the study, one plant was kept at an average daily temperature of 20 ∘C20~^\circ\text{C} but showed an unusually high weekly growth of 6565 mm. This data point was (20,65)(20, 65).

Comment on whether the horticulturalist should leave this data point in their analysis. Explain how this data point would affect the correlation.

[3]

Question 4

EasyPaper 1 · calculator5 marks
(a)

A marine biologist is studying the relationship between ocean temperature and the growth rate of a specific coral species. After collecting data, she calculates Pearson's product moment correlation coefficient (rr) for several pairs of variables. For each of the following rr values, choose from the list to accurately describe the correlation: perfect positive; strong positive; weak positive; zero; weak negative; strong negative; perfect negative.

(a) r=0.91r = 0.91

[1]
(b)

(b) r=−1r = -1

[1]
(c)

(c) r=0.23r = 0.23

[1]
(d)

(d) r=−0.78r = -0.78

[1]
(e)

(e) r=0r = 0

[1]

Question 5

MediumPaper 1 · calculator5 marks
(a)

The following table shows the number of days, dd, since a new mobile application was launched and the percentage of its initial active user base, xx, remaining at the beginning of that day.

Days since launch (dd)259141822
Percentage of active users left (xx)854832242017

The following table shows the natural logarithm of both dd and xx on these days, rounded to 2 decimal places.

ln⁡(d)\ln (d)0.691.612.202.642.893.09
ln⁡(x)\ln (x)4.443.873.473.183.002.83

Use the data in the second table to find the value of mm and the value of bb for the regression line, ln⁡x=m(ln⁡d)+b\ln x = m(\ln d) + b.

[2]
(b)

Assuming that the model found in part (a) remains valid, estimate the percentage of active users remaining when d=25d = 25.

[3]

Question 6

HardPaper 2 · calculator21 marks
(a)

Dr. Anya Sharma, a sports scientist, is investigating the relationship between training habits and performance in junior athletes. She wants to collect data on the weekly training hours of junior swimmers. She decides to interview every 5th swimmer entering the training facility until she has a sample of 50 swimmers.

State the sampling method Dr. Sharma has used.

[1]
(b)

Dr. Sharma constructed the following box and whisker diagram to show the weekly training hours (in hours) of a sample of junior swimmers.

A box and whisker diagram showing weekly training hours. The minimum is 2, the first quartile (Q1) is 4, the median is 6, the third quartile (Q3) is 9, and the maximum is 12.

Write down the median weekly training hours.

[1]
(c)

Calculate the interquartile range for the weekly training hours.

[2]
(d)

One swimmer in the sample reported training for 15 hours per week. Dr. Sharma believes this swimmer's training time is not an outlier.

Determine whether Dr. Sharma is correct. Support your reasoning.

[4]
(e)

Dr. Sharma also collected data on the average weekly training hours (xx) and the competition score (yy) for a group of athletes. These data are represented on the scatter diagram.

A scatter diagram showing competition score (y-axis from 0 to 120) versus weekly training hours (x-axis from 0 to 25). The points show a general negative correlation, with data points roughly between 5 and 20 hours.

Describe the correlation between weekly training hours and competition score.

[1]
(f)

Dr. Sharma correctly calculates the equation of the regression line yy on xx for these athletes to be y=−2.5x+110y = -2.5x + 110. She uses the equation to estimate the competition score for an athlete who trains 3 hours per week.

Find the competition score calculated by Dr. Sharma.

[2]
(g)

State whether it is valid to use the regression line yy on xx for Dr. Sharma's estimate in part (f). Give a reason for your answer, assuming the original data for training hours ranged from 5 to 20 hours.

[2]
(h)

Dr. Sharma investigated the relationship between an athlete's national competition rank and their average daily protein intake (in grams). She collected data for eight athletes, as shown in the table.

AthleteABCDEFGH
Competition Rank (RcompR_{comp})12345678
Protein Intake (g) (PintakeP_{intake})180150200160140190170130

Dr. Sharma intends to analyse the data using Spearman's rank correlation coefficient, rsr_s.

Copy and complete the information in the following table.

AthleteABCDEFGH
Rank - Competition Rank1
Rank - Protein Intake
[2]
(i)(i)

Calculate the value of rsr_s.

[3]
(i)(ii)

Interpret your result.

[3]

Question 7

MediumPaper 1 · calculator11 marks
(a)(i)

A digital marketing agency is analyzing the relationship between weekly advertising spend and sales revenue for a new product. The following table shows the advertising spend (in thousands of dollars) and the corresponding sales revenue (in thousands of dollars) over five weeks.

Advertising Spend SS (in $1000s) | 2 | 4 | 6 | 8 | 10

--|---|---|---|---|---

Sales Revenue RR (in $1000s) | 25 | 33 | 41 | 49 | 57

Find the equation of the regression line of RR on SS.

[4]
(a)(ii)

Describe the correlation between RR and SS with reference to the value of rr, the Pearson's product-moment correlation coefficient.

[3]
(b)

The marketing agency plans to spend $7500 on advertising next week. Estimate the sales revenue for the product.

[2]
(c)

The agency is considering a special campaign with an advertising spend of $50000. Explain why it would be inappropriate to use the equation found in part (a) to estimate the sales revenue for this campaign.

[2]

Question 8

HardPaper 2 · calculator14 marks
(a)

(a) An architect is designing a decorative archway for a park entrance. The cross-section of one half of the archway is modeled. The archway is symmetrical about the y-axis.

The architect models the base section of the archway as a straight line passing through the points (0,2)(0, 2) and (2,4)(2, 4), where all units are in metres.

Find the equation of the line passing through these two points.

[2]
(b)(i)

(b) The architect initially models the curved upper section of the archway using the following measured points:

(2,4)(2, 4), (4,5)(4, 5), (5.5,3)(5.5, 3), and (7,0)(7, 0).

(i) Find the equation of the least squares regression quadratic curve for these four points.

[2]
(b)(ii)

(ii) By considering the gradient of this curve when x=2x = 2, explain why it may not be a good model for the archway.

[1]
(c)

(c) The architect decides that a better model for the curved section would be a quadratic curve with a maximum point at (4.5,5.5)(4.5, 5.5) and that passes through the endpoint (7,0)(7, 0).

Find the equation of this new quadratic model.

[4]
(d)(i)

(d) Believing this to be a better model for the archway, the architect wants to estimate the volume of the solid generated by rotating this half-archway about the x-axis.

(i) Write down an expression for this estimate of the volume as a sum of two integrals.

[4]
(d)(ii)

(ii) Find the value of this estimate.

[1]

Question 9

MediumPaper 1 · calculator6 marks
(a)

Two renowned film critics, Alice and Bob, independently rank the artistic merit of ten independent films. The films are labelled F1 to F10, and their rankings are shown in the table below.

FilmAlice's RankBob's Rank
F112
F221
F334
F443
F555
F667
F776
F889
F998
F101010

Write down the rank that Alice awards film F4.

[1]
(b)

Calculate Spearman's rank correlation coefficient for these data.

[4]
(c)

Comment on your answer to part (b) in terms of the ranks awarded by Alice and Bob.

[1]

Question 10

HardPaper 2 · calculator17 marks
(a)

Emily and Liam are researching the adoption of smart home devices in a specific region to create a model predicting future usage. They collect the following data:

YearYears after 2000 (xx)Number of devices (in thousands) (NN)
2000010
20055150
201010400
201515750
2020201000
2025251100

Emily proposes the number of devices can be modelled using quadratic regression to find a function of the form N(x)=ax2+bx+cN(x) = ax^2 + bx + c, where xx is the number of years after 2000.

Find the equation of Emily's model.

[3]
(b)

Emily finds the coefficient of determination for her model is 0.979770.97977 to five significant figures.

State whether the coefficient of determination supports Emily's proposal. Justify your answer.

[2]
(c)

Comment on the validity of Emily's model with reference to one of the parameters in the equation.

[1]
(d)

(i) Find the value of N′(25)N'(25) and interpret this value in context.

(ii) By considering the changes in device adoption in the table, use the value found in part (d)(i) to comment on the validity of Emily's model.

[4]
(e)

Liam proposes that the device adoption instead follows a logistic model of the form

G(x)=15001+149e−0.15xG(x) = \frac{1500}{1+149 \text{e}^{-0.15x}}

where xx is the number of years after 2000 and G(x)G(x) is the number of devices in thousands.

State a reason why it may be valid to use Liam's proposal to predict future device adoption.

[1]
(f)

(i) Find G′(x)G'(x).

(ii) Hence find the year, according to Liam's model, during which the greatest device adoption growth rate occurred.

[6]

Question 11

MediumPaper 1 · calculator11 marks
(a)(i)

A manufacturing company is investigating the relationship between a worker's years of experience and their daily production output. They randomly selected five workers and recorded their years of experience (xx) and their average daily production output (yy, in units).

The data is shown in the following table:

Years of Experience (xx)Daily Production Output (yy)
1106
3114
5126
7134
9147

The company believes there might be a linear relationship between a worker's experience and their production output.

(a)(i) Find the Pearson's product moment correlation coefficient, rr, for this data.

[4]
(a)(ii)

(a)(ii) Find the equation of the least squares regression line of yy on xx for this data.

[4]
(b)

(b) According to this model, calculate how many more units a worker would produce daily if they had an extra 2.5 years of experience.

[2]
(c)

(c) State one reason why the value obtained in part (b) might not be valid for all workers.

[1]

Question 12

HardPaper 3 · calculator24 marks
(a)(i)

Ms. Anya Sharma, a school principal, wants to investigate if the number of hours students spend studying affects their exam scores. This question asks you to review Ms. Sharma's methods and conclusions.

Ms. Sharma obtained a list of students from her school. She contacted them and asked them to fill in an anonymous questionnaire. Participants were asked to state their weekly study hours and their most recent exam score (out of 100). Of the 250 students on the list, 11 replied.

Ms. Sharma's results are summarized in the following table:

Student IDWeekly Study Hours (X)Exam Score (Y)
1550
2765
3860
41078
51270
6655
7972
81180
9445
101385
111860

Describe one way in which Ms. Sharma could improve the reliability of her investigation.

[1]
(a)(ii)

Describe one criticism that can be made about the validity of Ms. Sharma's investigation.

[1]
(b)

Ms. Sharma classifies Student 11 as an outlier and removes their data from the analysis. Suggest one possible justification for her decision to remove it.

[1]
(c)(i)

For the remaining ten student responses in the table, Ms. Sharma calculates the mean exam score to be 6666. Calculate the mean weekly study hours for these remaining responses.

[2]
(c)(ii)

Determine the value of rr, Pearson's product-moment correlation coefficient, for these remaining responses.

[2]
(d)(i)

Ms. Sharma decides to carry out a hypothesis test on the correlation coefficient to investigate whether increased weekly study hours are associated with higher exam scores. State why the hypothesis test should be one-tailed.

[1]
(d)(ii)

State the null and alternative hypotheses for this test.

[2]
(d)(iii)

The critical value for this test, at the 5% significance level, is 0.549. Ms. Sharma assumes that the population is bivariate normal. Determine whether there is significant evidence of a positive correlation between weekly study hours and exam scores. Justify your answer.

[2]
(e)(i)

Ms. Sharma wants to create a model to predict how changing weekly study hours might affect exam scores. To do this, she assumes that weekly study hours, XX, is the independent variable and the exam score, YY, is the dependent variable.

She first considers a linear model of the form Y=aX+bY = aX + b. Use Ms. Sharma's data to find the value of aa and of bb.

[1]
(e)(ii)

Interpret, referring to study hours and exam scores, what the value of aa represents.

[1]
(e)(iii)

Ms. Sharma then considers a quadratic model of the form Y=cX2+dX+eY = cX^2 + dX + e. Find the value of cc, of dd and of ee.

[1]
(e)(iv)

Find the coefficient of determination for each of the two models she considers.

[2]
(e)(v)

Hence compare the two models.

[1]
(e)(vi)

Ms. Sharma decides to use the coefficient of determination to choose between these two models. Comment on the validity of her decision.

[1]
(f)(i)

After presenting the results of her investigation, a colleague questions whether Ms. Sharma's sample is representative of all students in the school. A report states that the mean weekly study hours for all students in the school is 99 hours. Ms. Sharma decides to carry out a test to determine whether her sample could realistically be taken from a population with a mean of 99 hours. State the name of the test which Ms. Sharma should use.

[1]
(f)(ii)

State the null and alternative hypotheses for this test.

[1]
(f)(iii)

Perform the test, using a 5% significance level, and state your conclusion in context.

[3]

Question 13

MediumPaper 1 · calculator9 marks
(a)

Brew & Bloom, a local coffee shop, decided to investigate if there was any correlation between their daily coffee sales and the average daily hours of sunlight. They collected the following data over 12 days:

Day | Hours of Sunlight (SS) | Coffee Sales (CC, in units)

---|---|---

1 | 3.23.2 | 240240

2 | 4.54.5 | 260260

3 | 5.85.8 | 308308

4 | 6.16.1 | 333333

5 | 7.37.3 | 328328

6 | 8.08.0 | 345345

7 | 8.58.5 | 394394

8 | 9.19.1 | 393393

9 | 7.07.0 | 316316

10 | 5.55.5 | 298298

11 | 4.04.0 | 241241

12 | 3.53.5 | 228228

(a) Plot the given points on a scatter diagram.

[2]
(b)

(b) Calculate the coordinates of the mean point (Sˉ,Cˉ)(\bar{S}, \bar{C}) and hence draw a line of best fit on your scatter diagram from part (a).

[3]
(c)

(c) Calculate Pearson's Product Moment Correlation Coefficient for this data, and interpret your result.

[2]
(d)

(d) Comment on whether the coffee shop owner can conclude, from this data, that daily hours of sunlight directly affect coffee sales.

[2]

Question 14

HardPaper 3 · calculator28 marks
(a)(i)

(a) TechInnovate is considering collecting more data for their analysis.

(i) State one advantage of increasing the sample size.

[1]
(a)(ii)

(ii) State one disadvantage of increasing the sample size.

[1]
(b)

(b) The production manager at Plant Alpha recorded the time, in minutes, taken to produce a batch of electronic components for 10 randomly selected batches:

18.2,19.5,17.8,20.1,18.5,19.0,17.5,20.5,18.8,19.318.2, 19.5, 17.8, 20.1, 18.5, 19.0, 17.5, 20.5, 18.8, 19.3

Find the value of sn−1s_{n-1} for this sample from Plant Alpha.

[2]
(c)

(c) A manager claims that Plant Alpha's production times are more consistent than Plant Beta's. Given that the sample standard deviation (sn−1s_{n-1}) for Plant Beta's production times is 1.051.05 minutes, make one criticism of this claim.

[1]
(d)(i)

(d) TechInnovate wants to compare the mean production times of Plant Alpha and Plant Beta using a pooled t-test.

(i) State the condition regarding population variances required to use a pooled t-test.

[1]
(d)(ii)

(ii) Given that for Plant Beta, a sample of 1212 batches yielded a mean production time of xˉB=19.3\bar{x}_B = 19.3 minutes and a sample standard deviation of sB=1.05s_B = 1.05 minutes, state whether TechInnovate should use a pooled t-test in this case. Justify your answer.

[2]
(e)(i)

(e) TechInnovate believes Plant Alpha has a lower mean production time than Plant Beta.

(i) State appropriate null and alternative hypotheses for the pooled t-test.

[2]
(e)(ii)

(ii) Find the p-value.

[2]
(e)(iii)

(iii) Given that the test is carried out at the 5% significance level, state the appropriate conclusion in context. Justify your answer.

[2]
(f)(i)

(f) The company also investigates the relationship between operator experience (in years) and the number of defective items produced per day. A sample of 8 operators yielded the following data:

Operator Experience (years)Number of Defective Items
215
510
313
87
118
69
412
78

(i) Assuming all requirements are met, perform a test at the 5% significance level to determine if there is a linear correlation between operator experience and the number of defective items. State the hypotheses and justify your conclusion.

[4]
(f)(ii)

(ii) If the requirements for this test are not met, state an alternative test that could be used.

[1]
(g)

(g) For the data in (f.i), the equation of the least squares regression line of defective items (DD) on operator experience (EE) is D=−1.5E+18.25D = -1.5E + 18.25. Give, in context, an interpretation of the gradient −1.5-1.5 in this model.

[1]
(h)(i)

(h) TechInnovate uses a baseline model to predict the number of defective items (DpredD_{pred}) for a batch based on its size (SS): Dpred=0.5S+10D_{pred} = 0.5S + 10. The "Quality Deviation" (QQ) for a batch is defined as Q=Dpred−DactualQ = D_{pred} - D_{actual}. A positive Quality Deviation indicates better-than-expected quality.

(i) Show that for a batch of 150150 components from Plant Beta that produced 8080 defective items, the Quality Deviation is 5.05.0.

[2]
(h)(ii)

(ii) To compare quality control, samples of Quality Deviation scores were collected:

  • Plant Alpha: nQA=15n_{QA} = 15, xˉQA=4.5\bar{x}_{QA} = 4.5, sQA=1.2s_{QA} = 1.2
  • Plant Beta: nQB=18n_{QB} = 18, xˉQB=3.8\bar{x}_{QB} = 3.8, sQB=1.1s_{QB} = 1.1

Assuming that the appropriate requirements are met, use a pooled t-test at a 5% significance level to determine if the mean Quality Deviation is higher in Plant Alpha than in Plant Beta. Write down your null and alternative hypotheses and justify your conclusion.

[4]
(i)

(i) Using the results from parts (e) and (h.ii), explain how each plant could claim they are performing better than the other plant.

[2]

Question 15

MediumPaper 2 · calculator7 marks
(a)

A small online retail company, 'GadgetHub', tracks its daily advertising expenditure and the corresponding number of units sold for a new product over a period of ten days.

Daily Advertising Spend (USD), xxUnits Sold, yy
551616
661919
772323
882525
992828
10103131
11113434
12123737
13134040
14144242

Draw a scatter graph to represent this information. Label the axes clearly.

[3]
(b)

Describe the correlation between daily advertising spend and units sold.

[2]
(c)

State whether you think the daily advertising spend has an effect on the number of units sold. Give a reason for your answer.

[2]

Question 16

HardPaper 3 · calculator27 marks
(a)

A botanical garden is researching the growth of a rare orchid species, Orchidaceae splendens, in a controlled greenhouse environment. They are investigating the long-term viability of maintaining a large population.

The population of Orchidaceae splendens, PP, measured in thousands, is shown in the following table. The time tt is measured in years from the beginning of 2010, where t∈Rt \in \mathbb{R}.

End of yearttPP (thousands)
201010.5
201120.8
201231.0
201341.4
201451.9
201562.6
201674.0

The garden models this data set using the logistic function P=2501+Ce−ktP = \frac{250}{1+Ce^{-kt}}, where C,k∈R+C, k \in \mathbb{R}^+.

In context, explain the significance of 250250 in this logistic function.

[1]
(b)(i)

Use the value of PP at t=1t = 1 to show that C=499ekC= 499e^k.

[2]
(b)(ii)

Use the value of PP at t=7t = 7 to find a second expression for CC.

[3]
(c)(i)

Use your answers to part (b) to find a value for kk.

[2]
(c)(ii)

Use your answers to part (b) to find a value for CC.

[1]
(d)

The botanical garden uses its model to predict values of PP at the end of the years 2011 to 2015. These values are shown in the following table, correct to one decimal place.

End of yearttPP (actual, thousands)Predicted values (thousands)
201120.80.7
201231.01.0
201341.4aa
201451.92.0
201562.62.8

Calculate the value of aa correct to one decimal place.

[2]
(e)

Using the value of aa correct to one decimal place, find the sum of square residuals, SSresSS_{res}, when using this model to predict the values of PP.

[2]
(f)

As a measure of how well the model fits the data, the botanical garden uses the error function, EE, where E=SSresnE = \sqrt{\frac{SS_{res}}{n}}, and nn is the number of predictions made using the model.

The garden decides they will use this model if EE is less than 0.150.15.

By finding the value of EE for the model, show that the garden will decide that the model can be used.

[3]
(g)

Use the model to predict the population of Orchidaceae splendens in the greenhouse at the end of 2025.

[2]
(h)

One of the challenges in maintaining the orchid population is providing sufficient specialized protective covers. The botanical garden estimates that 15%15\% of orchids will require a specialized protective cover, and each cover can protect 500500 individual orchids.

Use the model to find an expression for the total number of specialized protective covers required at time tt. Give your answer in thousands of covers.

[2]
(i)

At the end of 2014 (t=5t=5) there were 2.52.5 thousand specialized protective covers in the greenhouse, and at the end of 2016 (t=7t=7) there were 3.53.5 thousand.

Find the average number of specialized protective covers installed per year during this period.

[1]
(j)

The botanical garden assumes that the number of specialized protective covers will continue to increase linearly, at the same rate, until the population stabilizes.

Determine the value of tt, where t≥5t \ge 5, at which the number of specialized protective covers will first be insufficient to meet demand, and hence the year in which this occurs.

[6]

Question 17

MediumPaper 1 · calculator11 marks
(a)(i)

A market research firm collected paired, bivariate data (x,y)(x, y) on advertising spend and weekly product sales for a new item. The data showed a strong linear relationship, and the yy on xx line of best fit is given by y=mx+cy = mx + c.

When the advertising spend is 55 thousand dollars (x=5x=5), the estimated weekly sales are 1212 thousand units (y=12y=12).

When the advertising spend is 1010 thousand dollars (x=10x=10), the estimated weekly sales are 2222 thousand units (y=22y=22).

(a) Find the value of

(i) mm

[2]
(a)(ii)

(ii) cc

[3]
(b)

(b) State whether the correlation is positive or negative.

[1]
(c)

(c) Given that the advertising spend is 88 thousand dollars, find the estimated weekly sales.

[3]
(d)

(d) When advertising spend is 33 thousand dollars, find an estimate for the weekly sales.

[2]

Question 18

MediumPaper 2 · calculator12 marks
(a)(i)

A team of agricultural researchers is investigating the effectiveness of a new soil nutrient on crop yield. They conducted an experiment on 8 plots of land, varying the amount of nutrient applied (xx grams per square meter) and measuring the resulting crop yield (yy kg per square meter).

The data is given in the table below.

Nutrient amount (xx)Crop yield (yy)
1012.1
1514.8
2017.5
2519.9
3022.3
3525.0
4027.6
4530.2

(a) (i) Calculate the Pearson product moment correlation coefficient for this data, correct to three significant figures.

[2]
(a)(ii)

(ii) In two words, describe the linear correlation that is exhibited by this data.

[1]
(a)(iii)

(iii) Calculate the yy on xx line of best fit in the form y=ax+by = ax + b, giving the values of aa and bb correct to three significant figures.

[3]
(b)(i)

The research team conducted a follow-up experiment on 4 additional plots, which were located in a different region with slightly different soil conditions. This extra data is given in the table below.

Nutrient amount (xx)Crop yield (yy)
1210.5
2818.0
3820.0
4821.5

(b) (i) Calculate the Pearson product moment correlation coefficient for the combined data of all 12 plots, correct to three significant figures.

[2]
(b)(ii)

(ii) In two words, describe the linear correlation that is exhibited by the combined data.

[1]
(b)(iii)

(iii) Suggest a reason why it may not be valid to calculate the yy on xx line of best fit for the combined data.

[3]

Question 19

MediumPaper 2 · calculator7 marks
(a)

A student is investigating the relationship between the number of hours spent studying for a particular subject and the score achieved in the exam for that subject. They collected data from 10 classmates:

Hours Studied (hours) | Exam Score

---|---

2 | 54

3 | 55

4 | 66

5 | 67

6 | 68

7 | 74

8 | 81

9 | 78

10 | 86

11 | 94

Draw a scatter graph to represent this information.

[3]
(b)

Describe the correlation between the hours studied and the exam score.

[2]
(c)

Comment on whether the number of hours studied has an effect on the exam score.

[2]

Question 20

MediumPaper 1 · calculator12 marks
(a)(i)

A market research firm is studying the relationship between advertising expenditure (xx, in thousands of dollars) and monthly product sales (yy, in thousands of units) for a new gadget. They found that the paired bivariate data (x,y)(x, y) shows a strong linear correlation, modelled by the yy-on-xx line of best fit y=mx+cy = mx + c.

When the advertising expenditure is 55 thousand dollars, the estimated monthly sales are 2525 thousand units. When the advertising expenditure is 1212 thousand dollars, the estimated monthly sales are 4646 thousand units.

(a) Find the values of:

(i) mm

[3]
(a)(ii)

(ii) cc

[3]
(b)

(b) State whether there is positive or negative correlation.

[1]
(c)

(c) The market research firm plans an advertising expenditure of 88 thousand dollars. Find the estimated monthly sales, in thousands of units.

[3]
(d)

(d) When the advertising expenditure is 66 thousand dollars, find an estimate for the monthly sales, in thousands of units, given that this is an example of interpolation.

[2]

16 more Linear Correlation of bivariate data (scatter diagrams, lines of best fit, Pearson) + regression line interpretation questions in the app

Every answer is marked mark by mark, IB-style, and the AI tutor helps when you are stuck.

Where marks are lost

  • Rounding an intermediate value and then using it in a later part. Costs a mark every time, and AI's multi-part modelling questions give it more chances to happen than AA's shorter, more self-contained ones.
  • Using your own wrong value after failing a "show that." All follow through is withdrawn for the rest of that question.
  • Leaving an answer in calculator notation. Never accepted in a final answer, and AI's constant calculator use makes this the easiest slip in the whole subject.
Free. Every IB subject.
No card, no trial that runs out. Just a free account.
  • 50 marked answers a month
    Marked mark by mark, IB-style
  • Hints and mark schemes
    On every part of every question
  • 3,000+ questions
    All 6 subjects, SL and HL, mapped to the syllabus
  • Progress that adapts
    Your Study Profile picks what to practise next

Practise this topic as a session

Pick a difficulty and paper, and FourtyFive tracks your progress on this topic as you go.

or with email
FAQ

Questions,
answered.

Can't find what you're looking for? Email our student team.

What does Linear Correlation of bivariate data (scatter diagrams, lines of best fit, Pearson) + regression line interpretation cover in IB Maths AI?

Bivariate Data: Data on two variables, paired to examine relationships. Correlation vs. Causation: Correlation does not imply causation. Linear Correlation Types (Scatter Diagram):.

Is Linear Correlation of bivariate data (scatter diagrams, lines of best fit, Pearson) + regression line interpretation SL or HL?

Both. SL and HL students study Linear Correlation of bivariate data (scatter diagrams, lines of best fit, Pearson) + regression line interpretation to the same depth.

How do I revise Linear Correlation of bivariate data (scatter diagrams, lines of best fit, Pearson) + regression line interpretation for IB Maths AI?

Start from the core idea: bivariate Data: Data on two variables, paired to examine relationships. In the exam: one of the highest-frequency subtopics in the course. The pattern is: calculate r, describe the correlation in words, find the regression equation, use it to predict, then comment on the reliability of that prediction. Then practise exam-style questions, easiest first, writing out every step of your working before you check it.

How does FourtyFive help me practise Linear Correlation of bivariate data (scatter diagrams, lines of best fit, Pearson) + regression line interpretation?

FourtyFive has 36 Linear Correlation of bivariate data (scatter diagrams, lines of best fit, Pearson) + regression line interpretation questions. Every answer you write is marked mark by mark, IB-style, and you see where each mark was won or lost. Every part has a hint, the AI tutor helps you through the step you are stuck on, and your Study Profile picks what to practise next.

Is FourtyFive free for Linear Correlation of bivariate data (scatter diagrams, lines of best fit, Pearson) + regression line interpretation practice?

Yes. A free account gives you 50 marked answers a month, and you do not need a card to sign up.

Can I handwrite Linear Correlation of bivariate data (scatter diagrams, lines of best fit, Pearson) + regression line interpretation answers on an iPad?

Yes. In the FourtyFive iPad app you write your working by hand with Apple Pencil, the way you would on paper, and it is marked the same way.

Start with the IB question
bank built for you.

Free to start, no card needed. Thousands of syllabus-mapped questions, AI Examiner marking, your weakest topics first.