Skip to content
  1. IB Question Bank
  2. Maths AA
  3. Statistics & Probability
Topic 4.11 · SL and HL

Regression lines and reverse regression (x on y) + applications: notes and practice questions

Summary
  • Regression line (y on x): Predicts yy using xx, expressed as y=a+bxy = a + bx.
  • Reverse regression (x on y): Predicts xx using yy, expressed as x=c+dyx = c + dy.
  • Applications: Used to model relationships between two variables, predict outcomes, and assess correlation strength. Coefficient bb or dd represents the rate of change.
  • Ensure correct regression type for accurate predictions.

How it is examined

Describing the correlation wants two words, strength and direction, and one of them is usually dropped. The interpretation of aa and bb has to be in context with units. Extrapolation warnings are their own mark. Paper 2, 6 to 9 marks across parts. Short, and almost always attached to an SL 4.4 question as the last part. The marked idea is choosing the right line for the direction of the prediction. 2 to 4 marks, Paper 2.

Given in the booklet

Nothing. The AA booklet has no entry for Pearson's rr, no least-squares regression formula and no coefficient of determination. Every one of these comes from the GDC, which is why the guidance says technology should be used.

Key ideas
  • Linear correlation of bivariate data.
  • Pearson's product-moment correlation coefficient, rr.
  • Scatter diagrams; lines of best fit, by eye, passing through the mean point.
  • Equation of the regression line of yy on xx.

Linking questions

  • Aim 8: the correlation between smoking and lung cancer was "discovered" using mathematics, and science had to justify the cause.
  • The mirror of the SL 4.4 warning: yy on xx predicts yy, xx on yy predicts xx, and each is unreliable for the other direction.

Practice questions

23 questions · 2 easy · 19 medium · 2 hard
Showing 20 of 20

Question 1

EasyPaper 2 · calculator5 marks
(a)

Eight cars are tested to determine their fuel efficiency. Their weight, WW, in tonnes, and their fuel efficiency, FF, in kilometres per litre (km/L\text{km/L}), are shown in the table.

Weight (W tonnesW\text{ tonnes})1.201.351.501.101.651.401.251.55
Fuel efficiency (F km/LF\text{ km/L})18.516.214.020.112.515.817.613.2

The equation of the regression line of FF on WW for this data can be written in the form F=aW+bF = aW + b.

Find the value of aa and the value of bb.

[2]
(b)

Write down the value of the Pearson's product-moment correlation coefficient, rr.

[1]
(c)

Use the equation of the regression line of FF on WW to predict the fuel efficiency of a car with a weight of 1.45 tonnes1.45\text{ tonnes}.

[2]

Question 2

MediumPaper 2 · calculator6 marks
(a)

An agronomist studies the effect of monthly rainfall on the yield of a new variety of wheat. The monthly rainfall, RR in mm, and the wheat yield, YY in kg per hectare, were recorded for seven different regions.

The results are shown in the table below.

Monthly Rainfall, RR (mm)5075100120150180200
Wheat Yield, YY (kg/ha)410475560610690740800

The relationship between the variables can be modelled by the regression equation Y=aR+bY = aR + b.

(a) Find the value of aa and of bb.

[3]
(b)

(b) Write down the value of the Pearson's product-moment correlation coefficient, rr.

[1]
(c)

(c) Use the regression equation to estimate the wheat yield in a region where the monthly rainfall is 135 mm.

[2]

Question 3

HardPaper 2 · calculator19 marks
(a)(i)

(a) A study investigates the relationship between the amount of a specific fertilizer (xx kg) applied to a crop field and the resulting crop yield (yy tonnes). Data from 88 experimental plots is collected and presented in the table below.

xx (kg)yy (tonnes)
10105.55.5
12126.26.2
14147.07.0
16168.18.1
18189.09.0
20209.89.8
222210.510.5
242411.211.2

(i) Calculate the Pearson product moment correlation coefficient for this data.

[2]
(a)(ii)

(ii) In two words, describe the linear correlation that is exhibited by this data.

[1]
(a)(iii)

(iii) Calculate the yy on xx line of best fit, in the form y=ax+by = ax + b. Give the values of aa and bb to three significant figures.

[3]
(b)(i)

(b) Another four experimental plots are added to the study, with the following results:

xx (kg)yy (tonnes)
11119.09.0
15156.06.0
191912.012.0
23237.07.0

(i) Calculate the Pearson product moment correlation coefficient for the combined data of all 1212 plots.

[3]
(b)(ii)

(ii) In two words, describe the linear correlation that is exhibited by the combined data.

[1]
(b)(iii)

(iii) Suggest a reason why it would not be particularly valid to calculate the yy on xx line of best fit for the combined data.

[2]

Question 4

EasyPaper 1 · no calculator6 marks
(a)

A coffee shop owner records the average daily temperature, TT (in °C), and the number of hot coffees sold, CC, for a number of days. The scatter diagram shows the results.

Scatter diagram showing a negative correlation between temperature and coffee sales. The x-axis is Temperature (T) from 0 to 30. The y-axis is Number of hot coffees sold (C) from 100 to 300. Points are scattered generally from top-left to bottom-right.

The mean temperature for these days was 15 °C.

For these results, the equation of the regression line of CC on TT is C=−5T+250C = -5T + 250.

(a) Find the mean number of hot coffees sold.

[2]
(b)

(b) Draw the regression line on the scatter diagram.

[2]
(c)

(c) By placing a tick (✔) in the correct box, determine which of the following statements is true.

StatementCheckbox
The correlation is positive
The correlation is negative
There is no correlation
[1]
(d)

(d) Give a reason why the regression line should not be used to estimate the number of hot coffees sold when the average temperature is 35 °C.

[1]

Question 5

MediumPaper 1 · no calculator7 marks
(a)

A biologist is studying a species of fish. The length, LL cm, and weight, WW g, of each fish in a sample are recorded.

The lengths of the fish are summarized in the following box and whisker diagram.

Box and whisker diagram showing lengths of fish (cm) with minimum 5, Q1 12, median 15, Q3 18, maximum 30. all labelled

Find the largest value of LL that would not be considered an outlier.

[3]
(b)(i)

The regression line of WW on LL is W=25L−150W = 25L - 150. The regression line of LL on WW is L=0.03W+10.5L = 0.03W + 10.5.

One of the fish in the sample weighs 200 g. Estimate the length of this fish.

[2]
(b)(ii)

Find the mean weight of the fish in the sample.

[2]

Question 6

HardPaper 1 · no calculator14 marks
(a)(i)

A marine biologist is studying a species of sea turtle. She collects data on the carapace length, LL cm, and mass, MM kg, for 30 turtles. The Pearson's product-moment correlation coefficient for this data is found to be r=−0.92r = -0.92. The equation of the regression line of MM on LL is M=−2.5L+150M = -2.5L + 150.

The biologist discovers her measuring tape was misaligned, and all length measurements are 2 cm too short. Her weighing scale was also faulty, showing a mass 1.5 kg less than the true mass for each turtle. The data is corrected for these errors.

(i) State the new value of the Pearson's product-moment correlation coefficient, rr.

[1]
(a)(ii)

(ii) State the new value for the gradient of the regression line of MM on LL.

[1]
(a)(iii)

(iii) Briefly justify your answers to part (a)(i) and (a)(ii).

[2]
(b)(i)

The biologist decides to present her findings to an international conference and converts her original measurements to different units. She converts the original length measurements from cm to mm, and the original mass measurements from kg to g.

(i) State the new value of rr.

[1]
(b)(ii)

(ii) Find the new value for the gradient of the regression line of mass on length.

[2]
(b)(iii)

(iii) Briefly justify your answer for the new gradient.

[2]
(c)(i)

For a different analysis, the biologist defines a "size index", SS, as S=200−LS = 200 - L. She investigates the relationship between the size index SS and the original mass MM in kg.

(i) Find the value of rr for the correlation between SS and MM.

[2]
(c)(ii)

(ii) Find the gradient of the regression line of MM on SS.

[2]
(c)(iii)

(iii) Describe the linear correlation between the size index SS and the mass MM.

[1]

Question 7

MediumPaper 1 · no calculator7 marks
(a)

A study was conducted to investigate the relationship between the number of hours, hh, a student spends studying for an exam and their score, ss (%), in that exam.

The number of hours spent studying is summarized in the following box and whisker diagram.

Box and whisker diagram of hours spent studying, showing min=2, Q1=5, median=8, Q3=12, max=20

(a) Find the largest value of hh that would not be considered an outlier.

[3]
(b)(i)

The regression line of ss on hh is s=5h+25s = 5h + 25. The regression line of hh on ss is h=0.1s+2.5h = 0.1s + 2.5.

(b) (i) One of the students scored 90% on the exam. Estimate the number of hours they studied.

[2]
(b)(ii)

(ii) Find the mean score of all the students in the study.

[2]

Question 8

MediumPaper 2 · calculator7 marks
(a)(i)

A tutor wants to investigate the relationship between the number of hours a student spends studying for a mathematics test and the score they achieve on the test. They collect data from five students:

Number of hours studied (xx)25748
Test score (yy)5570856590

The relationship between xx and yy can be modelled by the regression line of yy on xx with equation y=ax+by = ax + b.

Find the value of aa and the value of bb.

[3]
(a)(ii)

Write down the value of Pearson's product-moment correlation coefficient, rr.

[1]
(b)

Interpret, in context, the value of aa found in part (a)(i).

[1]
(c)

Another student studies for 6 hours for the mathematics test.

Use the regression line from part (a)(i) to estimate this student's test score.

[2]

Question 9

MediumPaper 2 · calculator5 marks
(a)

A university lecturer is investigating the relationship between the number of hours, HH, students spend studying for a particular module each week and their final exam score, SS, out of 120. The results for eight randomly selected students are summarized in the table below.

Study Hours (HH )5781012141517
Exam Score (SS )60687582889598105

(a) Find Pearson's product-moment correlation coefficient, rr, for these data.

[2]
(b)

(b) The relationship between the variables can be modelled by the regression equation S=aH+bS = aH + b. Write down the value of aa and the value of bb.

[1]
(c)

(c) One student, who currently studies 10 hours per week, decides to increase their study time by an extra three hours per week. Based on the given data, determine by how many marks their final exam score could be expected to change.

[2]

Question 10

MediumPaper 2 · calculator7 marks
(a)

A tech company, "InnovateTech", is investigating the relationship between the average weekly training hours of its software developers and their quarterly productivity scores (out of 150). A sample of eight developers' data is collected and summarized in the table below.

Average weekly training hours (h)Productivity Score (P)
1282
1894
25116
30133
1586
22104
35145
28124

Find Pearson's product-moment correlation coefficient, rr, for these data.

[2]
(b)

The relationship between the variables can be modelled by the regression equation P=ah+bP = ah + b. Write down the value of aa and the value of bb.

[1]
(c)

InnovateTech is considering providing an optional advanced training module. Based on the given data, determine how a developer's productivity score could be expected to alter if they completed this module, which adds an extra five hours of training per week.

[2]
(d)

The CEO of InnovateTech asserts that increased training hours directly cause higher productivity scores. Comment on the validity of the CEO's assertion.

[1]
(e)

InnovateTech later discovered that due to a data entry error, all recorded productivity scores were exactly 10 points lower than their true values. The data was corrected by adding 10 points to each developer's productivity score.

State how, if at all, the value of rr would be affected.

[1]

Question 11

MediumPaper 2 · calculator7 marks
(a)

The total number of units produced, PP, by a factory depends on the number of hours, HH, the factory operates. A production manager uses the model P=−0.8H2+28H+50P = -0.8H^2 + 28H + 50 to predict the total units produced on any given day, where 5≤H≤205 \le H \le 20.

An energy auditor investigates the relationship between the total units produced and the energy consumption, EE, in kilowatt-hours (kWh). The following table shows the data collected on five different days.

Units Produced (P)Energy Consumption (E, in kWh)22516.225017.427518.829019.530520.3\begin{array}{|c|c|} \hline \textbf{Units Produced (P)} & \textbf{Energy Consumption (E, in kWh)} \\ \hline 225 & 16.2 \\ 250 & 17.4 \\ 275 & 18.8 \\ 290 & 19.5 \\ 305 & 20.3 \\ \hline \end{array}

Use the production model to estimate the number of units produced when the factory operates for 15 hours.

[2]
(b)

Find an appropriate regression equation that will allow the auditor to predict the energy consumption on a day when PP units are produced.

[3]
(c)

Hence, use your regression equation to predict the energy consumption when the factory operates for 15 hours.

[2]

Question 12

MediumPaper 2 · calculator4 marks
(a)

A marketing analyst is investigating the relationship between the amount spent on online advertising and the number of product units sold.

The following table shows the advertising spend, xx (in thousands of dollars), and the corresponding number of units sold, yy (in hundreds), for a new product over seven different campaigns.

Advertising Spend (xx, in thousands of dollars)13710151820
Units Sold (yy, in hundreds)5585150215310360395

The value of Pearson's product-moment correlation coefficient, rr, for this data is 0.9990.999, correct to three significant figures.

The regression line of yy on xx for this data can be written in the form y=ax+by = ax + b.

Find the value of aa and the value of bb.

[2]
(b)

Use your regression line to estimate the number of units sold when the advertising spend is 1212 thousand dollars.

[2]

Question 13

MediumPaper 2 · calculator8 marks
(a)

A botanist is investigating the effect of a new plant nutrient on the yield of a specific fruit crop. They apply varying amounts of the nutrient to several plants and record the amount of nutrient applied (xx grams per plant) and the resulting fruit yield (yy kg per plant). The results are shown in the table below.

Amount of nutrient (xx grams)Fruit yield (yy kg)
10101.31.3
20201.91.9
30302.42.4
40403.13.1
50503.33.3
60604.24.2

The botanist wants to model the relationship between the amount of nutrient and the fruit yield using a linear regression. Find the equation of the regression line of yy on xx.

[3]
(b)

Use your equation from part (a) to estimate the fruit yield for a plant that received 3535 grams of the nutrient.

[2]
(c)

Write down the correlation coefficient rr.

[1]
(d)

Using the value of rr, describe the correlation between the amount of nutrient and the fruit yield.

[2]

Question 14

MediumPaper 2 · calculator11 marks
(a)(i)

A marketing firm analyzes the relationship between advertising expenditure and monthly product sales. They find that the sales, SS (in thousands of units), are strongly correlated with the advertising expenditure, AA (in thousands of dollars). The line of best fit for SS on AA is of the form S=mA+cS = mA + c.

When the advertising expenditure is A=8A = 8 thousand dollars, the estimated sales are S=120S = 120 thousand units.

When the advertising expenditure is A=15A = 15 thousand dollars, the estimated sales are S=232S = 232 thousand units.

Find the values of:

i mm

[2]
(a)(ii)

ii cc.

[3]
(b)

State if there is positive or negative correlation.

[1]
(c)

The marketing firm plans an advertising expenditure of A=12A = 12 thousand dollars. Find the estimated monthly sales, SS.

[3]
(d)

When the advertising expenditure is A=10A = 10 thousand dollars, find an estimate for the value of SS, given that this is an interpolation.

[2]

Question 15

MediumPaper 2 · calculator16 marks
(a)

A teacher wants to investigate the relationship between students' performance in Mathematics and Physics. They collect the final exam scores (out of 100) for 10 randomly selected students. The bivariate data obtained is given in the table below.

Student | 11 | 22 | 33 | 44 | 55 | 66 | 77 | 88 | 99 | 1010

---|---|---|---|---|---|---|---|---|---|---

Math score (xx) | 7272 | 8080 | 6868 | 7575 | 9090 | 8383 | 7070 | 9595 | 6262 | 8888

Physics score (yy) | 7575 | 8383 | 7070 | 7878 | 9292 | 8585 | 7373 | 9898 | 6565 | 9090

(a) Find the Pearson product moment correlation coefficient, rr, for this data.

[2]
(b)

(b) State, in two words, a description for this linear correlation.

[2]
(c)(i)

(c) Find the equation of the line of best fit for:

i The Physics score (yy) on the Math score (xx).

[3]
(c)(ii)

ii The Math score (xx) on the Physics score (yy).

[3]
(d)

(d) Another student scored 8686 in Mathematics but was unable to take the Physics exam. Estimate the score they would have obtained in Physics, giving your answer to the nearest integer.

[2]
(e)

(e) A different student scored 7777 in Physics but did not take the Mathematics exam. Estimate the score they would have obtained in Mathematics, giving your answer to the nearest integer.

[2]
(f)

(f) If a student scored 110110 in Mathematics, explain why it would be unreliable to use a line of best fit to estimate their Physics score.

[2]

Question 16

MediumPaper 1 · no calculator16 marks
(a)(i)

A researcher investigates the relationship between the number of hours spent studying per week, hh, and the score on a standardized test, ss, for a group of 40 students. The Pearson product-moment correlation coefficient is found to be r=0.75r = 0.75, and the line of best fit for ss on hh is given by s=3.2h+40s = 3.2h + 40.

The researcher decides to adjust the data. The number of study hours for each student is reduced by 1, and the test score for each student is increased by 5.

(a) (i) State the new value of rr for the adjusted data.

[1]
(a)(ii)

(ii) State the new value for the gradient of the ss on hh line of best fit.

[1]
(a)(iii)

(iii) Give a reason for your answers to (i) and (ii).

[2]
(a)(iv)

(iv) Describe in two words the linear correlation that exists for this new data.

[1]
(b)(i)

The original data, with r=0.75r = 0.75 and s=3.2h+40s = 3.2h + 40, is now converted to different units. The study hours are converted from hours to minutes, and the test scores are scaled by a factor of 1.5.

(b) (i) State the new value of rr.

[1]
(b)(ii)

(ii) Find the new value for the gradient of the line of best fit of the scaled scores on the scaled hours.

[2]
(b)(iii)

(iii) Give a reason for your answers to (i) and (ii).

[2]
(c)(i)

Using the original data again (r=0.75r = 0.75 and s=3.2h+40s = 3.2h + 40), a new variable called 'study deficit', dd, is defined as d=20−hd = 20 - h. The relationship between ss and dd is investigated.

(c) (i) State the new value of rr for ss and dd.

[1]
(c)(ii)

(ii) Find the new line of best fit for ss on dd.

[2]
(c)(iii)

(iii) Give a reason for your answers to (i) and (ii).

[2]
(c)(iv)

(iv) Describe in two words the linear correlation that exists for the new data.

[1]

Question 17

MediumPaper 2 · calculator9 marks
(a)

A fitness researcher collected data on the average number of hours an individual spends exercising per week (xx) and their average weekly calorie expenditure (yy, in hundreds of calories). The results for a sample of individuals are shown in the table below.

Hours Exercised (xx)Calories Burned (yy, in hundreds)
2.02.05.55.5
3.53.58.08.0
5.05.011.011.0
6.56.514.514.5
8.08.017.017.0

This data can be modelled by the regression line with equation y=ax+by = ax + b.

Write down the values of aa and of bb.

[2]
(b)

Explain what the gradient, aa, represents in this context.

[2]
(c)

Use the model to estimate the average weekly calories burned if an individual exercises for 4.54.5 hours per week.

[3]
(d)

Explain why it would be unreliable to use this model to predict the calories burned for someone exercising 1515 hours per week.

[2]

Question 18

MediumPaper 2 · calculator8 marks
(a)

A research team is studying the relationship between the average daily temperature (xx in ∘C^\circ C) during a growing season and the average height of a specific plant species (yy in cm) at harvest. They collected data over seven growing seasons.

Average Daily Temperature (xx in ∘C^\circ C)Plant Height (yy in cm)
18187272
20207878
22228585
24249090
26269696
2828102102
3030108108

(a) Write down the equation of the yy on xx regression line for this data, giving your coefficients to three significant figures.

[3]
(b)

(b) Estimate the average plant height if the average daily temperature during the growing season was 25∘C25^\circ C.

Give your answer to one decimal place.

[2]
(c)(i)

(c) The average daily temperatures were converted from Celsius to Fahrenheit using the formula F=95C+32F = \frac{9}{5}C + 32.

For each of the following quantities, state whether it would change or remain the same:

(i) the mean of the average daily temperatures

[1]
(c)(ii)

(ii) the standard deviation of the average daily temperatures

[1]
(c)(iii)

(iii) the correlation coefficient, rr

[1]

Question 19

MediumPaper 2 · calculator8 marks
(a)

A local coffee shop records its weekly sales of two popular coffee beans, Espresso Blend and Single Origin, over 10 weeks. The sales figures, in kilograms (kg), are shown in the following table.

Weekly Sales (kg)

Cafe Week12345678910
Espresso Blend (xx)50657055807560908595
Single Origin (yy)57667467787471918194

The cafe manager determines that the equation of the regression line of yy on xx for these sales is y=0.712x+23.7y = 0.712x + 23.7.

Find the value of Pearson's product-moment correlation coefficient, rr.

[2]
(b)(i)

The new cafe manager uses the regression line of yy on xx (y=0.712x+23.7y = 0.712x + 23.7) for making predictions.

One week, Espresso Blend sales were 1515 kg. The manager estimated Single Origin sales to be 3434 kg. Give a reason why this estimation might not be appropriate.

[1]
(b)(ii)

Another week, Single Origin sales were 8585 kg. The manager used the line y=0.712x+23.7y = 0.712x + 23.7 to estimate Espresso Blend sales as 8686 kg. Give a reason why this method is not appropriate.

[1]
(c)

Use an appropriate method to show that the estimated Espresso Blend sales for the week when Single Origin sales were 8585 kg is 8585 kg (to the nearest integer).

[4]

Question 20

MediumPaper 2 · calculator4 marks
(a)

A group of six students recorded the number of hours they spent studying for a mathematics exam, xx, and their corresponding score on the exam, yy, out of 100.

The data is presented in the table below:

Weekly Study Hours (xx)Exam Score (yy)
860
1068
1275
1482
1688
1895

The regression line of yy on xx for this data can be written in the form y=ax+by = ax + b.

Find the value of aa and the value of bb. Give your answers to three significant figures.

[2]
(b)

Use the equation of the regression line to estimate the exam score of a student who studies for 1515 hours per week.

[2]

3 more Regression lines and reverse regression (x on y) + applications questions in the app

Every answer is marked mark by mark, IB-style, and the AI tutor helps when you are stuck.

Where marks are lost

  • Rounding an intermediate value and then using it.
  • Answering to the wrong accuracy. Two significant figures, or six, where the rule says exactly or three.
  • Writing the answer and nothing else.
Free. Every IB subject.
No card, no trial that runs out. Just a free account.
  • 50 marked answers a month
    Marked mark by mark, IB-style
  • Hints and mark schemes
    On every part of every question
  • 3,000+ questions
    All 6 subjects, SL and HL, mapped to the syllabus
  • Progress that adapts
    Your Study Profile picks what to practise next

Practise this topic as a session

Pick a difficulty and paper, and FourtyFive tracks your progress on this topic as you go.

or with email
FAQ

Questions,
answered.

Can't find what you're looking for? Email our student team.

What does Regression lines and reverse regression (x on y) + applications cover in IB Maths AA?

Regression line (y on x): Predicts y using x, expressed as y = a + bx. Reverse regression (x on y): Predicts x using y, expressed as x = c + dy. Applications: Used to model relationships between two variables, predict outcomes, and assess correlation strength. Coefficient b or d represents the rate of change.

Is Regression lines and reverse regression (x on y) + applications SL or HL?

Both. SL and HL students study Regression lines and reverse regression (x on y) + applications to the same depth.

How do I revise Regression lines and reverse regression (x on y) + applications for IB Maths AA?

Start from the core idea: regression line (y on x): Predicts y using x, expressed as y = a + bx. In the exam: describing the correlation wants two words, strength and direction, and one of them is usually dropped. The interpretation of a and b has to be in context with units. Then practise exam-style questions, easiest first, writing out every step of your working before you check it.

How does FourtyFive help me practise Regression lines and reverse regression (x on y) + applications?

FourtyFive has 23 Regression lines and reverse regression (x on y) + applications questions. Every answer you write is marked mark by mark, IB-style, and you see where each mark was won or lost. Every part has a hint, the AI tutor helps you through the step you are stuck on, and your Study Profile picks what to practise next.

Is FourtyFive free for Regression lines and reverse regression (x on y) + applications practice?

Yes. A free account gives you 50 marked answers a month, and you do not need a card to sign up.

Can I handwrite Regression lines and reverse regression (x on y) + applications answers on an iPad?

Yes. In the FourtyFive iPad app you write your working by hand with Apple Pencil, the way you would on paper, and it is marked the same way.

Start with the IB question
bank built for you.

Free to start, no card needed. Thousands of syllabus-mapped questions, AI Examiner marking, your weakest topics first.