Non-linear Regression, Evaluating least squares regression, sum of squares, and R^2: notes and practice questions
- Non-linear regression: Used when a curve fits bivariate data better than a straight line.
- Residual: Difference between actual and predicted value:
- Sum of Square Residuals () (HL only): Sum of squared residuals; smaller indicates a better model fit.
- Standard non-linear models:
- Linear:
- Quadratic:
- Cubic:
- Exponential: or
- Power:
- Sine:
- Sum of Square Residuals formula (HL only):
- Mean Squared Error (MSE):
- Linearising using logarithms (HL only):
- Power Model (): (linear between and ).
- Exponential Model (): (linear between and ).
- Finding non-linear regression equation: Enter data into GDC, select specified model, GDC calculates constants.
- Evaluating least squares regression (HL only): Calculate for each model; model with lowest is the best fit.
- Linearising data (HL only): Transform equation to linear logarithmic form, align with GDC's regression line (), solve for original constants.
- GDC tip: Plot scatter diagram and regression model graph for visual assessment of fit.
- HL topics: Non-linear regression, linearising using logarithms, and evaluating least squares regression curves using .
- Model selection: Exam question will specify the exact regression model to use.
- Extrapolation: Predictions made outside the original data range are unreliable.
How it is examined
The classic HL question fits two or three models to the same data and asks which is best, and the expected answer is not "the one with the highest ". The guide says so directly, so a full-mark answer weighs against the shape of the data and the context, including whether the model behaves sensibly outside the data range. is compared between models rather than computed from the formula.
- Regression with non-linear functions.
- The evaluation of least squares regression curves using technology.
- The sum of square residuals () as a measure of fit for a model.
- The coefficient of determination (), and its evaluation using technology.
Awareness that , and hence equals 1 if , may enhance understanding but will not be examined.
Linking questions
- Links to other subjects: evaluation of in graphical analysis (sciences).
Practice questions
1 question · 1 hardQuestion 1
HardPaper 3 · calculator24 marksMs. Anya Sharma, a school principal, wants to investigate if the number of hours students spend studying affects their exam scores. This question asks you to review Ms. Sharma's methods and conclusions.
Ms. Sharma obtained a list of students from her school. She contacted them and asked them to fill in an anonymous questionnaire. Participants were asked to state their weekly study hours and their most recent exam score (out of 100). Of the 250 students on the list, 11 replied.
Ms. Sharma's results are summarized in the following table:
| Student ID | Weekly Study Hours (X) | Exam Score (Y) |
|---|---|---|
| 1 | 5 | 50 |
| 2 | 7 | 65 |
| 3 | 8 | 60 |
| 4 | 10 | 78 |
| 5 | 12 | 70 |
| 6 | 6 | 55 |
| 7 | 9 | 72 |
| 8 | 11 | 80 |
| 9 | 4 | 45 |
| 10 | 13 | 85 |
| 11 | 18 | 60 |
Describe one way in which Ms. Sharma could improve the reliability of her investigation.
Describe one criticism that can be made about the validity of Ms. Sharma's investigation.
Ms. Sharma classifies Student 11 as an outlier and removes their data from the analysis. Suggest one possible justification for her decision to remove it.
For the remaining ten student responses in the table, Ms. Sharma calculates the mean exam score to be . Calculate the mean weekly study hours for these remaining responses.
Determine the value of , Pearson's product-moment correlation coefficient, for these remaining responses.
Ms. Sharma decides to carry out a hypothesis test on the correlation coefficient to investigate whether increased weekly study hours are associated with higher exam scores. State why the hypothesis test should be one-tailed.
State the null and alternative hypotheses for this test.
The critical value for this test, at the 5% significance level, is 0.549. Ms. Sharma assumes that the population is bivariate normal. Determine whether there is significant evidence of a positive correlation between weekly study hours and exam scores. Justify your answer.
Ms. Sharma wants to create a model to predict how changing weekly study hours might affect exam scores. To do this, she assumes that weekly study hours, , is the independent variable and the exam score, , is the dependent variable.
She first considers a linear model of the form . Use Ms. Sharma's data to find the value of and of .
Interpret, referring to study hours and exam scores, what the value of represents.
Ms. Sharma then considers a quadratic model of the form . Find the value of , of and of .
Find the coefficient of determination for each of the two models she considers.
Hence compare the two models.
Ms. Sharma decides to use the coefficient of determination to choose between these two models. Comment on the validity of her decision.
After presenting the results of her investigation, a colleague questions whether Ms. Sharma's sample is representative of all students in the school. A report states that the mean weekly study hours for all students in the school is hours. Ms. Sharma decides to carry out a test to determine whether her sample could realistically be taken from a population with a mean of hours. State the name of the test which Ms. Sharma should use.
State the null and alternative hypotheses for this test.
Perform the test, using a 5% significance level, and state your conclusion in context.
Reliability concerns the consistency and repeatability of the results. How can she ensure her measurements are more consistent or less prone to random error?
Validity concerns whether the study measures what it intends to measure and whether the results are generalizable. Are there other factors influencing exam scores? Is "study hours" accurately measured?
Look at the data for Student 11 compared to the general trend. What makes it unusual?
Sum the weekly study hours for the remaining students and divide by .
Use your GDC's statistical functions to calculate Pearson's for the data points (excluding Student 11).
Consider the specific direction of the relationship Ms. Sharma is investigating.
The null hypothesis typically states no effect or no relationship, while the alternative hypothesis states the effect or relationship you are looking for. Use the correct symbol for population correlation.
Compare your calculated value from part (c.ii) with the given critical value.
Use your GDC's linear regression function (LinReg(ax+b) ) with the data points.
The coefficient in a linear model represents the change in for every one-unit increase in .
Use your GDC's quadratic regression function (QuadReg) with the data points.
The coefficient of determination, , is often provided by your GDC along with the regression equation. For the linear model, .
A higher value generally indicates a better fit for the data.
tends to increase with the number of independent variables or parameters in a model, even if the additional terms do not significantly improve the model's predictive power.
This is a test comparing a sample mean to a known population mean when the population standard deviation is unknown (which is usually the case).
The null hypothesis assumes the sample comes from the population with the stated mean. The alternative hypothesis states it does not.
Use your GDC's t-test function (T-Test) for one sample. Input the sample data (study hours), the hypothesized population mean, and the significance level.
No question on this page matches those filters. Try another difficulty or paper.
Every Non-linear Regression, Evaluating least squares regression, sum of squares, and R^2 question, marked for you
Every answer is marked mark by mark, IB-style, and the AI tutor helps when you are stuck.
Where marks are lost
- Answering to the wrong accuracy. Two significant figures, or six, where the rule says exactly or three. Common wherever a GDC's full decimal display gets copied straight down.
- Rounding an intermediate value and then using it in a later part. Costs a mark every time, and AI's multi-part modelling questions give it more chances to happen than AA's shorter, more self-contained ones.
- Writing the answer and nothing else, where the mark scheme has an explicit M1 rather than an implied one. A bare answer cannot score full marks there.