Linear Correlation of bivariate data (scatter diagrams, lines of best fit, Pearson): notes and practice questions
- Scatter diagrams display pairs of data points to show the relationship between two variables.
- The line of best fit (or regression line) represents the linear relationship, typically found using the least squares method.
- Pearson’s correlation coefficient (r) quantifies the strength and direction of a linear relationship, with values between -1 and 1:
- indicates a perfect positive correlation,
- indicates a perfect negative correlation,
- indicates no linear correlation.
How it is examined
Describing the correlation wants two words, strength and direction, and one of them is usually dropped. The interpretation of and has to be in context with units. Extrapolation warnings are their own mark. Paper 2, 6 to 9 marks across parts.
Nothing. The AA booklet has no entry for Pearson's , no least-squares regression formula and no coefficient of determination. Every one of these comes from the GDC, which is why the guidance says technology should be used.
- Linear correlation of bivariate data.
- Pearson's product-moment correlation coefficient, .
- Scatter diagrams; lines of best fit, by eye, passing through the mean point.
- Equation of the regression line of on .
Linking questions
- Aim 8: the correlation between smoking and lung cancer was "discovered" using mathematics, and science had to justify the cause.
Practice questions
40 questions · 4 easy · 32 medium · 4 hardQuestion 1
EasyPaper 1 · no calculator5 marksAn IB student is conducting five different investigations for their mathematics internal assessment. For each investigation, they collect bivariate data and calculate the Pearson product-moment correlation coefficient, .
Use the list below to state the correct description for the following values of .
perfect positive, strong positive, weak positive, zero, weak negative, strong negative, perfect negative
(a) For an investigation into daily temperature and the number of people at a public pool, .
(b) For an investigation into a car's engine size and its fuel efficiency, .
(c) For an investigation into the number of hours a person trains per week and their time to complete a marathon, .
(d) For an investigation into a person's shoe size and their monthly phone bill, .
(e) For an investigation into the side length of a square and its perimeter, .
Consider the strength and direction of the correlation. A value close to 1 or -1 is strong. A value close to 0 is weak or zero. The sign indicates the direction (positive or negative).
The value is negative, so the correlation is negative. Is the absolute value of closer to 0 or 1?
The value is negative. Compare its absolute value to the thresholds for 'weak' and 'strong' correlation.
This value is very close to 0. What does that imply about the relationship between the variables?
What does an value of exactly 1 signify about the linear relationship between the two variables?
Question 2
MediumPaper 2 · calculator6 marksAn agronomist studies the effect of monthly rainfall on the yield of a new variety of wheat. The monthly rainfall, in mm, and the wheat yield, in kg per hectare, were recorded for seven different regions.
The results are shown in the table below.
| Monthly Rainfall, (mm) | 50 | 75 | 100 | 120 | 150 | 180 | 200 |
|---|---|---|---|---|---|---|---|
| Wheat Yield, (kg/ha) | 410 | 475 | 560 | 610 | 690 | 740 | 800 |
The relationship between the variables can be modelled by the regression equation .
(a) Find the value of and of .
(b) Write down the value of the Pearson's product-moment correlation coefficient, .
(c) Use the regression equation to estimate the wheat yield in a region where the monthly rainfall is 135 mm.
Enter the two lists of data (Rainfall and Yield) into your GDC and use the linear regression (LinReg y=ax+b) function to find the coefficients of the line of best fit.
The value for 'r' is calculated by your GDC at the same time as the regression coefficients. If you don't see it, you may need to turn on 'DiagnosticOn' in your calculator's settings.
Substitute the value R = 135 into the equation you found in part (a).
Question 3
HardPaper 2 · calculator19 marks(a) A study investigates the relationship between the amount of a specific fertilizer ( kg) applied to a crop field and the resulting crop yield ( tonnes). Data from experimental plots is collected and presented in the table below.
| (kg) | (tonnes) |
|---|---|
(i) Calculate the Pearson product moment correlation coefficient for this data.
(ii) In two words, describe the linear correlation that is exhibited by this data.
(iii) Calculate the on line of best fit, in the form . Give the values of and to three significant figures.
(b) Another four experimental plots are added to the study, with the following results:
| (kg) | (tonnes) |
|---|---|
(i) Calculate the Pearson product moment correlation coefficient for the combined data of all plots.
(ii) In two words, describe the linear correlation that is exhibited by the combined data.
(iii) Suggest a reason why it would not be particularly valid to calculate the on line of best fit for the combined data.
Use your GDC's statistical functions to calculate the Pearson product moment correlation coefficient (). Ensure you input the and values correctly into two lists.
Consider the value of the correlation coefficient you calculated in part (a)(i). What does a value close to 1 indicate about the relationship between and ?
Use your GDC's linear regression function (e.g., LinReg(ax+b) ). Make sure to specify as the dependent variable and as the independent variable.
Combine all values and all values into two new lists. Then, use your GDC to calculate the Pearson product moment correlation coefficient for this new, larger dataset.
How does the new correlation coefficient compare to the one calculated in part (a)(i)? What does this new value suggest about the linear relationship?
Consider the value of the correlation coefficient for the combined data. What does a low value indicate about the suitability of a linear model? Also, think about how the new data points might visually affect the overall trend.
Question 4
EasyPaper 1 · no calculator3 marksA biologist is studying the relationship between the concentration of a nutrient, (in mg/L), and the weekly growth rate of a particular plant, (in cm/week). The relationship is found to be linear. The regression line of on passes through the mean point and has a gradient of .
Estimate the weekly growth rate of a plant when the nutrient concentration is mg/L.
The regression line is a straight line. First, find the equation of this line using the given point and gradient. Then, substitute the given concentration value into your equation.
Question 5
MediumPaper 2 · calculator7 marksThe following table shows the advertising spending (in thousands of dollars) and the corresponding monthly sales (in thousands of units) for a new product over several months.
| Advertising Spending (x) | Monthly Sales (y) |
|---|---|
| 5 | 13 |
| 8 | 15 |
| 10 | 17 |
| 12 | 19 |
| 15 | 21 |
| 18 | 23 |
| 20 | 24 |
| 22 | 25 |
The data is also represented on the following scatter diagram.

The relationship between advertising spending (x) and monthly sales (y) can be modelled by the regression line of y on x with equation , where .
Write down the value of and the value of .
Use this model to predict the monthly sales (in thousands of units) when the advertising spending is 25 thousand dollars.
Write down the value of and the value of .
Draw the line of best fit on the scatter diagram.
Use your GDC to find the linear regression equation . Input the x-values and y-values into your calculator's statistics function.
Substitute the given value of x into the regression equation you found in part (a).
Use your GDC to find the mean of the x-values and the mean of the y-values from the given data.
Remember that the line of best fit must pass through the mean point and have the correct slope (determined by 'a').
Question 6
HardPaper 2 · calculator12 marksA marketing analyst is investigating the relationship between the amount spent on social media advertising, (in hundreds of dollars), and the monthly sales revenue, (in thousands of dollars), for a new product. The following data was collected over eight months:
| (hundreds of dollars) | ||||||||
|---|---|---|---|---|---|---|---|---|
| (thousands of dollars) |
The analyst models the data using the on regression line with equation .
Write down the values of and of .
Explain what the gradient represents in the context of this problem.
Explain what the -intercept represents in the context of this problem.
Estimate the monthly sales revenue if the company spends $350 on social media advertising.
The company aims to achieve a monthly sales revenue of $40000. Find how much they should spend on social media advertising, according to this model.
Explain why it might be inappropriate to use this model to predict the monthly sales revenue if the company spends $10000 on social media advertising.
Use your GDC's linear regression function (e.g., LinReg(ax+b) ) to find the values of and from the given data.
The gradient represents the change in the dependent variable () for every unit increase in the independent variable (). Consider the units of and .
The -intercept is the value of the dependent variable () when the independent variable () is zero. Consider what means in this scenario.
Substitute the given advertising spend into the regression equation you found in part (a). Remember that is in hundreds of dollars.
Set to the target sales revenue and solve for using the regression equation. Remember that is in thousands of dollars.
Consider the range of the original data for . What happens when you try to predict values far outside this range?
Question 7
EasyPaper 2 · calculator5 marksEight cars are tested to determine their fuel efficiency. Their weight, , in tonnes, and their fuel efficiency, , in kilometres per litre (), are shown in the table.
| Weight () | 1.20 | 1.35 | 1.50 | 1.10 | 1.65 | 1.40 | 1.25 | 1.55 |
|---|---|---|---|---|---|---|---|---|
| Fuel efficiency () | 18.5 | 16.2 | 14.0 | 20.1 | 12.5 | 15.8 | 17.6 | 13.2 |
The equation of the regression line of on for this data can be written in the form .
Find the value of and the value of .
Write down the value of the Pearson's product-moment correlation coefficient, .
Use the equation of the regression line of on to predict the fuel efficiency of a car with a weight of .
Enter the data into your GDC's statistics mode to find the linear regression coefficients.
Look for the value in the same GDC screen where you found and .
Substitute into your regression equation .
Question 8
MediumPaper 2 · calculator7 marksA tutor wants to investigate the relationship between the number of hours a student spends studying for a mathematics test and the score they achieve on the test. They collect data from five students:
| Number of hours studied () | 2 | 5 | 7 | 4 | 8 |
|---|---|---|---|---|---|
| Test score () | 55 | 70 | 85 | 65 | 90 |
The relationship between and can be modelled by the regression line of on with equation .
Find the value of and the value of .
Write down the value of Pearson's product-moment correlation coefficient, .
Interpret, in context, the value of found in part (a)(i).
Another student studies for 6 hours for the mathematics test.
Use the regression line from part (a)(i) to estimate this student's test score.
Use your GDC's statistics function to perform linear regression (LinReg(ax+b) ). Input the values into List 1 and the values into List 2.
The Pearson's product-moment correlation coefficient () is usually calculated by your GDC at the same time as the regression line.
The value of represents the change in for every unit increase in . Consider what and represent in this problem.
Substitute the given number of hours () into the regression equation you found in part (a)(i).
Question 9
HardPaper 1 · no calculator14 marksA marine biologist is studying a species of sea turtle. She collects data on the carapace length, cm, and mass, kg, for 30 turtles. The Pearson's product-moment correlation coefficient for this data is found to be . The equation of the regression line of on is .
The biologist discovers her measuring tape was misaligned, and all length measurements are 2 cm too short. Her weighing scale was also faulty, showing a mass 1.5 kg less than the true mass for each turtle. The data is corrected for these errors.
(i) State the new value of the Pearson's product-moment correlation coefficient, .
(ii) State the new value for the gradient of the regression line of on .
(iii) Briefly justify your answers to part (a)(i) and (a)(ii).
The biologist decides to present her findings to an international conference and converts her original measurements to different units. She converts the original length measurements from cm to mm, and the original mass measurements from kg to g.
(i) State the new value of .
(ii) Find the new value for the gradient of the regression line of mass on length.
(iii) Briefly justify your answer for the new gradient.
For a different analysis, the biologist defines a "size index", , as . She investigates the relationship between the size index and the original mass in kg.
(i) Find the value of for the correlation between and .
(ii) Find the gradient of the regression line of on .
(iii) Describe the linear correlation between the size index and the mass .
The data is being corrected by adding a constant value to all length measurements and another constant value to all mass measurements. How does such a transformation (a translation) affect the correlation coefficient?
The gradient of the regression line is given by . How does adding a constant to all data points affect the standard deviations and ?
Consider the definitions of correlation and standard deviation. A translation shifts the entire data cloud without changing its shape, spread, or orientation.
The conversion from cm to mm and kg to g involves multiplying the data by positive constants. How does scaling by a positive constant affect the correlation coefficient?
The new measurements are and . The gradient is affected by the scaling of both variables. Use the formula .
Explain how the scaling of each variable affects their respective standard deviations and, consequently, the gradient of the regression line using the formula .
The new variable is . This is a linear transformation of . How does multiplying a variable by a negative number affect the correlation coefficient?
The gradient is . You have the new from part (c)(i). How does the transformation affect the standard deviation of the x-variable?
The description should include both the strength and the direction of the correlation, based on the value of you found in (c)(i).
Question 10
EasyPaper 1 · no calculator6 marksA coffee shop owner records the average daily temperature, (in °C), and the number of hot coffees sold, , for a number of days. The scatter diagram shows the results.

The mean temperature for these days was 15 °C.
For these results, the equation of the regression line of on is .
(a) Find the mean number of hot coffees sold.
(b) Draw the regression line on the scatter diagram.
(c) By placing a tick (✔) in the correct box, determine which of the following statements is true.
| Statement | Checkbox |
|---|---|
| The correlation is positive | |
| The correlation is negative | |
| There is no correlation |
(d) Give a reason why the regression line should not be used to estimate the number of hot coffees sold when the average temperature is 35 °C.
The regression line always passes through the point of mean values, . You are given the mean temperature, , and the equation of the line.
To draw a straight line, you need two points. You found one point in part (a), which is the point of means. The equation of the line can give you another point, for example, the y-intercept.
Observe the general trend of the data points on the scatter diagram. As the temperature increases, what happens to the number of coffees sold?
Compare the value of 35 °C to the range of temperatures for which data was collected, as shown on the scatter diagram.
Question 11
MediumPaper 2 · calculator7 marksA botanist is studying the growth of a particular plant species. They record the average height of several plants (in cm) at different weeks after planting. The data collected is shown in the table below.
| Week (x) | Height (y) (cm) |
|---|---|
| 2 | 8.4 |
| 4 | 10.9 |
| 6 | 14.5 |
| 8 | 18.2 |
| 10 | 19.8 |
| 12 | 22.8 |
The relationship between the week number (x) and the plant height (y) can be modelled by the regression line of y on x with equation , where .
Write down the value of and the value of .
Use this model to predict the height of a plant after 15 weeks.
Write down the mean week number, , and the mean plant height, .
Draw the line of best fit on a scatter diagram for this data.

Use your GDC to perform linear regression on the given data. Input the week numbers into one list and the corresponding heights into another. The calculator will provide the values for 'a' (slope) and 'b' (y-intercept) for the regression line . Remember to round to an appropriate number of significant figures, usually three.
Substitute the given week number (x = 15) into the regression equation that you found in part (a). Make sure to use the unrounded values for 'a' and 'b' from your GDC for the most accurate result before rounding the final answer.
Your GDC can calculate the mean of the x-values and the mean of the y-values when you perform linear regression or use basic statistical functions. Look for and in the statistical output.
The line of best fit must pass through the mean point that you found in part (c). Use the slope 'a' and y-intercept 'b' from part (a) to accurately draw the line. Plot at least two points (e.g., the y-intercept and the mean point) and connect them with a straight line using a ruler.
Question 12
HardPaper 2 · calculator23 marksThe following table shows the annual revenue of a tech startup, Quantum Innovations, years after its launch in 2015.
| (years after 2015) | 0 | 2 | 4 | 6 | 8 |
|---|---|---|---|---|---|
| (revenue in millions of USD) | 1.5 | 2.8 | 4.2 | 5.5 | 7.1 |
A data analyst uses linear regression to model the revenue of Quantum Innovations using these data.
The analyst's model is .
(a)(i) Write down the value of and the value of .
(a)(ii) Interpret, in context, the value of .
(b) The analyst uses this model to predict the revenue of Quantum Innovations in the year 2030, where , and calculates a revenue of approximately million USD.
Comment on the reliability of the analyst's prediction.
(c)(i) A financial expert, Elena, develops an exponential model for Quantum Innovations' future revenue.
In this model, represents the revenue in millions of USD years after 2015, where .
Use Elena's model to predict the revenue of Quantum Innovations in the year 2035.
(c)(ii) Interpret, in context, the value in Elena's model.
(d) Another financial expert, Carlos, develops a third model for Quantum Innovations' revenue.
In this model, represents the revenue in millions of USD years after 2015, where .
Use Carlos's model to predict the revenue of Quantum Innovations in the year 2035.
(e) Determine the year in which the difference between the predictions from Elena's model and Carlos's model is greatest.
(f)(i) Find the value of
;
(f)(ii) Find the value of
.
(g) Compare and interpret, in context, the values of and .
Use your GDC's linear regression (LinReg(ax+b) ) function to find the values of and . Ensure you input the values as your independent variable and values as your dependent variable.
The value of represents the slope of the linear model. Think about what the slope means in terms of the variables (revenue) and (years).
Consider the range of the original data used to create the model. What happens when you make a prediction outside this range?
First, determine the value of that corresponds to the year 2035. Then substitute this value into Elena's exponential model.
In an exponential growth model , the base represents the growth factor. How is this related to a percentage growth rate?
As in part (c.i), first find the correct value of for the year 2035. Then substitute it into Carlos's model.
Define a difference function, for example, . Use your GDC to graph this function over the domain and find its maximum value. Remember to convert back to a year.
Use your GDC's numerical derivative function (e.g., nDeriv or dy/dx) to evaluate the derivative of Elena's model at . Alternatively, find the analytical derivative of and substitute .
Similar to part (f.i), use your GDC's numerical derivative function or find the analytical derivative of Carlos's model and substitute .
The derivative represents the instantaneous rate of change. Compare which model predicts a faster rate of revenue increase at and explain what that means for Quantum Innovations.
Question 13
MediumPaper 2 · calculator5 marksA university lecturer is investigating the relationship between the number of hours, , students spend studying for a particular module each week and their final exam score, , out of 120. The results for eight randomly selected students are summarized in the table below.
| Study Hours ( ) | 5 | 7 | 8 | 10 | 12 | 14 | 15 | 17 |
|---|---|---|---|---|---|---|---|---|
| Exam Score ( ) | 60 | 68 | 75 | 82 | 88 | 95 | 98 | 105 |
(a) Find Pearson's product-moment correlation coefficient, , for these data.
(b) The relationship between the variables can be modelled by the regression equation . Write down the value of and the value of .
(c) One student, who currently studies 10 hours per week, decides to increase their study time by an extra three hours per week. Based on the given data, determine by how many marks their final exam score could be expected to change.
Use your GDC's statistics functions to calculate Pearson's . Input the study hours as your independent variable and exam scores as your dependent variable.
Use your GDC's linear regression function (e.g., LinReg(ax+b) or LinReg(a+bx) ) to find the values of and . Pay attention to which variable is the slope and which is the y-intercept.
The coefficient in the regression equation represents the change in for every one-unit increase in .
Question 14
MediumPaper 2 · calculator7 marksA tech company, "InnovateTech", is investigating the relationship between the average weekly training hours of its software developers and their quarterly productivity scores (out of 150). A sample of eight developers' data is collected and summarized in the table below.
| Average weekly training hours (h) | Productivity Score (P) |
|---|---|
| 12 | 82 |
| 18 | 94 |
| 25 | 116 |
| 30 | 133 |
| 15 | 86 |
| 22 | 104 |
| 35 | 145 |
| 28 | 124 |
Find Pearson's product-moment correlation coefficient, , for these data.
The relationship between the variables can be modelled by the regression equation . Write down the value of and the value of .
InnovateTech is considering providing an optional advanced training module. Based on the given data, determine how a developer's productivity score could be expected to alter if they completed this module, which adds an extra five hours of training per week.
The CEO of InnovateTech asserts that increased training hours directly cause higher productivity scores. Comment on the validity of the CEO's assertion.
InnovateTech later discovered that due to a data entry error, all recorded productivity scores were exactly 10 points lower than their true values. The data was corrected by adding 10 points to each developer's productivity score.
State how, if at all, the value of would be affected.
Use your GDC to enter the data into two lists and calculate the linear regression statistics. Pearson's is one of the outputs.
The values for and are also outputs from your GDC's linear regression calculation. Remember to round to 3 significant figures.
Consider what the coefficient in the regression equation represents. How does a change in affect ?
Think about the fundamental difference between correlation and causation in statistics. Does a strong correlation automatically imply one variable causes the other?
Consider how adding a constant to all values of one variable affects the spread and relative positions of the data points, and thus the correlation coefficient.
Question 15
MediumPaper 2 · calculator6 marksA rare vintage comic book, 'The Cosmic Crusader #1', was valued at $8000 on January 1st 2015. Its value is projected to increase by 3.5% on January 1st each year.
Find the projected value of 'The Cosmic Crusader #1' for the year 2025, to the nearest dollar.
Another rare comic book, 'Galactic Guardian #1', has had its value tracked over several years. The values for various years are shown in the following table.
| Year (x) | Annual Value () |
|---|---|
| 2015 | 8000 |
| 2017 | 9050 |
| 2019 | 9980 |
| 2021 | 11020 |
| 2023 | 11950 |
Assuming 'Galactic Guardian #1''s annual value can be approximately modelled by the equation , use your GDC to show that 'Galactic Guardian #1' is projected to have a higher value than 'The Cosmic Crusader #1' in the year 2024, according to the model.
Remember that the value increases each year. Identify the initial value, the annual growth rate, and the number of growth periods from the starting year to the target year.
Use your GDC's linear regression function (e.g., LinReg(ax+b) ) to find the values of 'a' and 'b'. Then, substitute the target year into the equation to find the predicted value. Don't forget to compare it to the value of 'The Cosmic Crusader #1' in the same year.
Question 16
MediumPaper 2 · calculator7 marksThe total number of units produced, , by a factory depends on the number of hours, , the factory operates. A production manager uses the model to predict the total units produced on any given day, where .
An energy auditor investigates the relationship between the total units produced and the energy consumption, , in kilowatt-hours (kWh). The following table shows the data collected on five different days.
Use the production model to estimate the number of units produced when the factory operates for 15 hours.
Find an appropriate regression equation that will allow the auditor to predict the energy consumption on a day when units are produced.
Hence, use your regression equation to predict the energy consumption when the factory operates for 15 hours.
Substitute the given number of hours into the quadratic model for production.
Determine which variable is the independent variable () and which is the dependent variable () for the regression. Use your GDC to find the linear regression equation.
Take the number of units produced from part (a) and substitute it into the regression equation found in part (b).
Question 17
MediumPaper 2 · calculator5 marksA group of students recorded the number of hours they spent studying for a mathematics exam, , and their corresponding exam score, . The data is shown in the table below.
| Hours studying () | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|
| Exam score () | 58 | 65 | 71 | 79 | 86 | 91 | 96 |
The regression line of on for this data can be written in the form .
Find the value of and the value of .
Write down the value of the Pearson's product-moment correlation coefficient, .
Use your regression line to estimate the exam score of a student who studied for 9 hours.
Use your GDC's linear regression function (e.g., LinReg(ax+b) or a+bx) to find the slope () and y-intercept () for the given data.
The Pearson's product-moment correlation coefficient () is typically calculated by your GDC along with the regression line.
Substitute the given number of hours into your regression equation from part (a) to estimate the score.
Question 18
MediumPaper 2 · calculator4 marksA marketing analyst is investigating the relationship between the amount spent on online advertising and the number of product units sold.
The following table shows the advertising spend, (in thousands of dollars), and the corresponding number of units sold, (in hundreds), for a new product over seven different campaigns.
| Advertising Spend (, in thousands of dollars) | 1 | 3 | 7 | 10 | 15 | 18 | 20 |
|---|---|---|---|---|---|---|---|
| Units Sold (, in hundreds) | 55 | 85 | 150 | 215 | 310 | 360 | 395 |
The value of Pearson's product-moment correlation coefficient, , for this data is , correct to three significant figures.
The regression line of on for this data can be written in the form .
Find the value of and the value of .
Use your regression line to estimate the number of units sold when the advertising spend is thousand dollars.
Use your GDC to perform linear regression on the given data. Input the advertising spend as your independent variable () and the units sold as your dependent variable (). The calculator will provide the values for and for the regression line . Remember to round your answers to three significant figures.
Substitute the given advertising spend value into the regression equation you found in part (a). Make sure to use the unrounded values of and for the calculation, and then round your final answer to an appropriate number of significant figures.
Question 19
MediumPaper 2 · calculator9 marks(a) A mathematics teacher investigates the relationship between the number of hours a student spends studying for a test and their final score on the test. The teacher collected data from five students, as shown in the table below:
| Hours Studied (x) | Test Score (y) |
|---|---|
| 2 | 62 |
| 3 | 70 |
| 4 | 78 |
| 5 | 85 |
| 6 | 91 |
The relationship between the hours studied, x, and the test score, y, can be modelled by the regression line of y on x with equation .
(i) Find the value of a and the value of b.
(ii) Write down the value of Pearson's product-moment correlation coefficient, r.
(b) Interpret, in context, the value of a found in part (a)(i).
(c) On another occasion, a student studied for 7 hours for the test.
Use the regression line from part (a)(i) to estimate this student's test score.
Use your GDC to perform a linear regression (y on x) on the given data. Input the x-values into one list and the y-values into another.
The Pearson's product-moment correlation coefficient (r) is typically calculated by the GDC when performing linear regression. Look for the 'r' value in the regression output.
The value of 'a' represents the slope of the regression line. Consider what a positive or negative slope means in the context of hours studied and test scores.
Substitute the given number of hours studied into the regression equation that you found in part (a)(i).
Question 20
MediumPaper 2 · calculator7 marks(a) The expected crop yield, , in kilograms (kg), from a certain field depends on the amount of fertilizer, , applied in kg. A farmer models the relationship using the equation , where .
Use this model to estimate the crop yield when the farmer applies 18 kg of fertilizer.
(b) The farmer also investigates the relationship between the crop yield, , and the total profit, , in dollars ($). The following table shows the data collected from five different harvests.
| Crop Yield ( in kg) | Total Profit ( in $) |
|---|---|
| 339 | 140 |
| 395 | 175 |
| 425 | 195 |
| 430 | 205 |
| 409 | 188 |
| Crop Yield ( in kg) | Total Profit ( in $) |
|---|---|
| 339 | 140 |
| 395 | 175 |
| 425 | 195 |
| 430 | 205 |
| 409 | 188 |
Find an appropriate regression equation that will allow the farmer to predict the total profit based on the crop yield.
(c) Hence, use your regression equation to predict the total profit when the farmer applies 18 kg of fertilizer.
Substitute the given value of fertilizer amount into the provided quadratic model for the crop yield.
You need to find the linear regression equation where Profit (P) is the dependent variable and Crop Yield (Y) is the independent variable. Use your GDC's linear regression function (e.g., 'a + bx' or 'ax + b').
First, recall the crop yield you calculated in part (a) for 18 kg of fertilizer. Then, substitute this yield value into the regression equation you found in part (b) to predict the profit.
No question on this page matches those filters. Try another difficulty or paper.
20 more Linear Correlation of bivariate data (scatter diagrams, lines of best fit, Pearson) questions in the app
Every answer is marked mark by mark, IB-style, and the AI tutor helps when you are stuck.
Where marks are lost
- Using your own wrong value after failing a "show that".