Skip to main content
Advanced Statistical Methods

Regression Analysis: Significance Testing, Multiple Regression, and Categorical Predictors

Published: 2026-08-11
Level: postgraduate
Audience: Postgraduate students in Advanced Statistical Methods

Prerequisite Knowledge

This lecture builds on the following concepts from earlier lectures. If any feel unfamiliar, review the linked notes before proceeding.

Previously Covered in This Subject

  • Simple linear regression and the decomposition of variation (SST, SSR, SSE) — covered in Lecture 8
  • The coefficient of determination and reading regression output — covered in Lecture 8
  • Multiple linear regression, adjusted R-squared, and multicollinearity — covered in Lecture 8
  • T tests and F tests for regression significance — covered in Lecture 8
  • The one-way ANOVA table and degrees of freedom — covered in Lecture 8
  • The p-value approach and the one-sample T test — covered in Lecture 4
  • Null and alternative hypotheses, and one-tailed versus two-tailed tests — covered in Lectures 3 and 4
  • The normal distribution, Z-scores, and reading normal tables — covered in Lectures 2 and 3
  • The chi-square distribution — covered in Lecture 5

9.1 Mid-Semester Syllabus and Question Paper Expectations

9.1.1 Chapters, Sections, and What Is Omitted

Hook: You have two weeks of material to revise, a two-hour open-book paper, and one question designed to separate the top grades. The most valuable thing you can know right now is precisely what will not be asked — so you can spend your revision time only where it counts.

The mid-semester portion of the course is deliberately bounded. The syllabus focuses on the course textbook Chapters 7 and 8, and you are asked to stay inside the specific sections named for those chapters — do not wander outside them. Within these chapters, one group of sections is explicitly omitted: everything about population proportions. The reason is simple: the discussions for proportions mirror the population-mean discussions that were already covered, so the professor is not going to ask anything about proportions. The same applies to factorial experiments — nothing about them will appear. The stated chapters and sections are enough for the mid-semester; there is no need to put in any further effort beyond them.

Why is this a useful message rather than just an announcement? Because it tells you where the boundary of the course is. A chapter in a textbook is a long document; a syllabus is a contract. Everything inside the named sections is fair game, everything outside is a waste of time. Treat the syllabus as the map: chapters 7 and 8, specified sections only, no proportions, no factorial experiments.

There is one extra expectation attached to the syllabus: you should expect a few problems related to the normal distribution. Knowingly or unknowingly, the normal distribution is used everywhere in this portion — in hypothesis testing problems, in sampling distributions, and in the tests discussed below — so you are asked to revisit the normal distribution and, in particular, how the normal distribution tables are to be referred to. Nothing more than that is expected for this mid-semester.

The normal distribution is the workhorse underneath almost everything in this course: sampling distributions describe how a sample statistic such as behaves by leaning on the normal curve, and hypothesis tests (the Z test, and the T test and F test of Lecture 9's regression discussion) all compare a computed statistic against a known curve. "Revisit the normal distribution tables" is so not a topic request — it is a skills request: be able to look up a tail probability or a critical value quickly, because you will need that skill inside the hypothesis-testing questions.

A note on editions: the textbook may be the 12th or the 13th edition. The sections in the two editions are similar; only the problems may vary. Either edition is fine to study from.

9.1.2 Question Paper Structure

The mid-semester question paper contains around six or seven questions, planned for a duration of about two hours. The paper is a mixture of two kinds of questions: a few that you need to work out computationally, and a few that are output-related observation questions — you read a given output and comment on it. Instead of asking you to work everything out manually, part of the paper asks you to interpret results that are handed to you.

That division is important for how you practice. For the computational questions you must be fluent with the arithmetic: plugging numbers into the test statistics, computing sums of squares, reading the correct row of a table. For the output-observation questions you must be fluent in reading: the professor shows a regression output or an ANOVA table, and you have to state what the numbers mean — which coefficients are significant, what the p-values say, what conclusion the output supports. Both skills are exercised in this lecture's later sections, so this is not a general pep talk — the session is a training ground for exactly this paper format.

The paper is open book. Expect roughly half or one question that sits outside the routine — that question is the one designed to differentiate the grades. Everything asked needs to be addressed; it need not always be routine calculations.

Because the paper is open book, the exam is not a memory test — it is a decision test: given the tools, can you pick the right one and read the result correctly? The grade-differentiating question exists to separate students who memorized procedures from students who can apply judgment (for example, a question that looks like the routine ones but asks you to notice that a particular conclusion is not justified by the output).

The model paper that was shared is a replica of the same thing you can expect in the actual mid-semester paper. You cannot expect any drastic deviations from it. Working through the model paper gives a clear picture of how to address the questions and in what way to write your answers.

The model paper is your best rehearsal. Because the actual paper mirrors it, practice by answering the model paper under exam conditions — timed, on paper, without peeking — and then compare your answer structure (how you write the conclusion, how you present the numbers) as well as the numbers themselves. The professor's guidance is that presentation matters: the paper checks how you address a question, not merely that you reached a number.

9.1.3 Topics to Master for the Mid-Semester

The four things to be thorough with are the Z distribution, the chi-square distribution, the T distribution, and the ANOVA table — plus regression. These same topics are reflected in the model paper from earlier years of the same course. The advice in one line: be thorough with the Z distribution, chi-square, T, the ANOVA table, and the regressions. On difficulty: the paper is not as difficult as you may be anticipating; the expectation is that everyone should handle it comfortably.

Exam note: expect a few normal distribution problems, so practice reading normal distribution tables before the exam. Expect 6–7 questions in a 2-hour, open-book paper, mixing hand computations with "read the output and comment" questions, with a possible grade-differentiating question that is not a routine calculation.

What each of these five topics looks like at mid-semester level:

  • Z distribution — the standard normal curve: standardized values , reading tail areas from the normal table, and finding critical values .
  • Chi-square distribution — the right-skewed curve used for tests about a single population variance and for the tests of independence and goodness of fit; practice reading the chi-square table with degrees of freedom.
  • T distribution — the bell-shaped curve used when the population standard deviation is unknown and estimated from the sample; degrees of freedom matter for which row of the table you read.
  • ANOVA table — the layout of sources of variation (treatments, error, total) with sums of squares, degrees of freedom, mean squares, the F ratio, and the p-value; the regression ANOVA table used in this lecture is the same structure applied to regression.
  • Regression — estimating the regression equation, the coefficient of determination, significance testing with the T and F tests, and (after this lecture) multiple regression, adjusted , multicollinearity, and dummy variables.

Intuition: these five items are not five unrelated chapters — they are one chain. The Z and T distributions are the rulers you compare test statistics against; chi-square handles the variance-type tests; the ANOVA table organizes the sums of squares that feed the F test; regression is where the whole chain is used on a real modeling problem. If you study them in that order — curve, then table, then how the numbers enter the table — the model paper falls into place.

9.1.4 Student Questions and Answers

Q: Will there be questions outside the covered chapters, or things we have not discussed?

A: That is also possible — I cannot answer everything as to whether a given topic is there or not. The model paper is a replica of the mid-semester paper, so you can get clarity from it, and the paper is not so difficult that you should worry. Most of your queries about the mid-semester have been answered here; maximum clarity is already given.

The honest part of the answer is the first sentence: nobody can promise that every question sits inside the named sections. The practical part is the rest: since the model paper mirrors the real paper, working through it is the best available prediction of what will be asked. Note the reassurance built into the answer — the professor's expectation is that the paper is comfortably handled, not that it is a trap.

9.2 Simple Linear Regression Recap and the Coefficient of Determination

9.2.1 The Estimated Regression Equation

Hook: You are handed a regression output and asked: "Write the estimated regression equation." Most students read the whole table and hesitate. The answer is one line — and it is always the first thing you extract from the output.

The session starts by revisiting simple linear regression from the previous session. Simple linear regression — regression with exactly one dependent variable and one independent variable — was discussed using a specific data set, referred to here as the PJA problem, where the independent variable is the student population size and the dependent variable is the quarterly sales.

Recall: this is the Armand's Pizza Parlors style problem — 10 restaurant locations, each with a student population around it (in thousands) and its quarterly sales (in thousands of dollars). The question the model answers: can the size of the student population explain the ups and downs of quarterly sales?

From that data set, the estimated regression equation shown in the output is

where:

  • (read "y-hat") — the predicted value of quarterly sales for a given student population . The hat marks it as a prediction, not an observed value.
  • — the independent variable, the student population size.
  • — the intercept: the predicted value of when . It is where the line crosses the vertical axis.
  • — the slope: the predicted change in when increases by 1 unit.

For the PJA problem the numbers are and , giving

So if you are asked to write the estimated regression equation, you look at the coefficients column of the regression output and write exactly this form. This was the first thing we could extract from the output of the previous session.

Worked reading of the coefficients (PJA problem). The output's coefficients column shows two rows: the intercept (Constant) and the row for (student population). The Constant row holds ; the row holds . Writing them in the template gives . Interpretation: a campus with 1,000 more students is associated with predicted quarterly sales 5,000 dollars higher (because is in thousands of students and in thousands of dollars, each unit of adds 5 units of ).

Sense-check with a real point: for a campus of 10,000 students (), the prediction is , i.e., 110,000 dollars in quarterly sales — a plausible mid-range value for this data, whose observed sales run from 58,000 to 202,000 dollars. The equation passes the sanity test.

9.2.2 Coefficient of Determination

The second thing we can take from the output is how much of the total variation in the estimated model explains. That is the coefficient of determination, written , defined as the ratio of the regression sum of squares to the total sum of squares:

where:

  • — the sum of squares due to regression: the amount of total variation in that the estimated line explains.
  • — the total sum of squares: the total variation in the observed values around their mean.
  • — the fraction of total variation explained, always between 0 and 1 for the models here.

The professor's plain-language description: "the formula for is SSR by SST" — the amount of total variation explained by the estimated simple linear regression. There is a notation convention worth keeping straight: the lower-case is used for simple linear regression, while the capital notation is used for multiple linear regression.

Formalize — why the ratio means "fraction explained". Total variation measures how far each observed sits from the overall mean — think of it as the total scatter of the sales figures. When you fit a line, part of that scatter is captured: is how far the predicted values sit from the mean, i.e., how much scatter the model reproduces. The leftover scatter is , the distance of observed points from the fitted line. The identity (Section 9.2.3) then makes literally "the fraction of the scatter the model explains": an of 0.9 means nine-tenths of the wiggle in is reproduced by the line, one-tenth is left as noise.

For the PJA problem, the coefficient of determination equals . The verbal conclusion attached to this number: 90% of the variability in the sales can be explained by the relationship between the estimated regression — that is, between the student population and the sales.

Worked example: the PJA numbers behind . The previous session's output gives the ANOVA ingredients: , , so .

Sense-check: the result must be a fraction between 0 and 1 — it is. And means about 90% of the variation in quarterly sales is explained by the student-population line, leaving only about 10% unexplained. That is a strong fit for an economic data set.

9.2.3 The Three Variations

A graphical way of assessing the different types of variation was also introduced with this example: once you have a scatter plot and you fit the regression equation to it, there are three variations at play. The total variation (SST) splits into the variation due to the regression (SSR) and the variation due to chance, that is, the residual or error variation (SSE):

The three sum-of-squares quantities, side by side:

Quantity Name What it measures Formula
Total sum of squares Total scatter of around — "how much variation there is to explain"
Regression sum of squares Scatter of the predicted values around — "how much the line explains"
Error sum of squares Scatter of the observed points around the fitted line — "the leftover noise"

The three always satisfy : every piece of total variation is either captured by the line or left in the residuals. There is no third place for it to go.

Real-world: this decomposition is exactly what a retailer does when it asks "how much of the ups and downs in our quarterly sales can we explain with a single driver, and how much is left over as unexplained noise?"

9.2.4 Why Estimation Alone Is Not Enough

Pitfall — the high- trap: an of 0.9027 looks impressive, and it is tempting to stop there and declare the model good. The professor's warning is explicit: it is not enough to estimate the regression and look at the value. describes fit; it does not prove that the relationship is real. A line fitted to random noise can still show a respectable by luck — the sample size is small (only restaurants), so the number carries uncertainty that alone does not report.

The recap concludes with a motivation that frames the rest of the session: it is not enough to estimate the regression and look at the value. There is a need for further analysis to justify that the assumed model is the appropriate one. There are many ways to justify a model; the two we perform next are called the T test and the F test. The point being made: an of 0.9027 looks impressive, but you still need formal tests before you can recommend the estimated model as fine.

The mental picture to carry into Section 9.3: says "the line explains 90% of the scatter"; the T and F tests ask the sharper question "is the slope really non-zero, or could this happen by chance?" The first is a description of the sample; the second is a claim about the population. Both are needed before the model earns your trust.

9.3 Significance Testing of the Regression: T Test and F Test

9.3.1 The Hypotheses: Is the Slope Really Non-Zero?

Hook: Your estimated line explains 90% of the sales variation. But is the slope 5 real, or could a sample of ten restaurants produce a slope like that purely by chance? That is the question the T test and the F test answer.

The expected value of the dependent variable is written as a linear function of the independent variable:

where:

  • — the expected value (mean) of the dependent variable for a given . In the PJA problem, the mean quarterly sales for a campus with a given student population.
  • — the population intercept: the unknown true value that estimates.
  • — the population slope: the unknown true change in per unit of , estimated by .
  • — the independent variable, student population size.

The key logical move: if equals zero, then the model collapses to just — no is involved, which means no independent variable is influencing , and so and are not related to each other at all. The moment you justify that , you are saying X and Y are unrelated.

Why is the critical claim? Plug it into the model: . The term vanishes, so the mean of is the same constant regardless of the student population. If that were true, knowing the population size would tell you nothing about sales — the whole regression would be pointless. The regression is worth keeping only if .

The two hypotheses tested are then

  • (the null hypothesis) — "no linear relationship": the slope is zero.
  • (the alternative hypothesis) — "there is a linear relationship": the slope is non-zero.

The goal of the analysis is to reject the null hypothesis. If the null is rejected, takes some non-zero value and there exists a relationship between and . If we fail to reject, we cannot say anything about the relationship. So the entire significance testing exercise — both the T test and the F test — is a way to strengthen or verify the claim that the estimated regression equation is fine, and specifically that the relationship between and really exists.

Scope — what "fail to reject" does and does not mean: rejecting gives evidence that the relationship exists; failing to reject only says the sample evidence is not strong enough to prove it. It is not a proof that and it is not a reason to keep or discard the model by itself — it is a "no verdict" state. Also, significance is not causation: a significant slope means the variables are related, not that population size causes sales.

9.3.2 The T Test Statistic and Its Standard Error

The T test statistic mirrors the familiar shape of a standardized variable. Recall the standard normal variable

The professor's analogy: in the T test, the numerator is replaced by , the estimated slope, and the denominator is replaced by the sample standard error of the slope:

where:

  • — the estimated slope from the output (for PJA, ).
  • — the standard error of the slope: how much would wobble from sample to sample. It plays the role of in the z formula.
  • — the estimate of the error standard deviation , i.e., the typical size of a residual. For PJA, .
  • — the sum of squared deviations of the values around their mean , measuring how spread out the student populations are; the larger the spread of , the more precisely the slope can be estimated, and the smaller becomes.
  • — the number of observations (for PJA, ).

The professor's words: "I am giving the numerator values ... and the denominator... by square root of summation minus bar whole square." Connect this to the PJA data: is the quarterly sales and is the student population; with the given values, the denominator can be computed.

The structural similarity to is deliberate: both are "signal divided by noise". In the z formula, measures how far a value sits from its mean and is the spread of the distribution; here, measures the slope we found and is how much that slope would vary by chance. A slope that is large relative to its own uncertainty is a strong signal; a slope that is small relative to its uncertainty could easily be chance. The t statistic is that ratio.

The value of is not free-floating: is the square root of the mean squared error (MSE), which is itself an estimate of the error variance . The MSE is read straight off the ANOVA table.

9.3.3 The ANOVA Table and the Degrees of Freedom

The ANOVA table for the PJA regression has the standard three sources: regression, error (residual), and total. It reports, for each source, the degrees of freedom, the sum of squares, and the mean square. Here:

  • — the sum of squares due to the regression,
  • — the sum of squares due to the residual or error,
  • — the total sum of squares.

The error sum of squares for the PJA data is , based on observations. The error degrees of freedom is . Why ? Because the data set was used to estimate two parameters — and — so two degrees of freedom are deducted. The mean squared error is then computed as SSE over , that is, 1530 over 8:

and the standard error of estimate, , equals

Why the degrees of freedom work as they do. Each parameter estimated from the data spends one degree of freedom. Simple regression estimates two parameters ( and ), so of the observations only are "free" for measuring error. That is why the error row uses and why the mean square divides by it: is the average squared error per free observation, an unbiased estimate of the error variance . The complete PJA ANOVA table:

Source Sum of squares Degrees of freedom Mean square F p-value
Regression 14200 1 14200 74.25 .000
Error 1530 8 191.25
Total 15730 9

The degrees of freedom add up: , and the sums of squares add up: .

9.3.4 Worked Computation: The T Test for the PJA Data

The pieces are now assembled. The slope from the output is , the standard error of the slope is

and the T statistic comes out as

Worked example: every step of the T test for the PJA data.

Step 1 — the denominator spread. The ten student populations give . Its square root:

Step 2 — the standard error of the slope. Divide the standard error of estimate by that root:

Step 3 — the t statistic. Divide the slope by its standard error:

The professor noted in class "don't bother about how to calculate this 8.62" — the software prints it — but the arithmetic above is exactly what the software does, and knowing the route is what lets you check the output for consistency.

Step 4 — degrees of freedom. .

Step 5 — the verdict. Compare against the t table: with 8 degrees of freedom, already leaves area 0.005 in the upper tail. Our 8.62 is far beyond it, so the p-value is much smaller than 0.01 and the null is rejected.

Sense-check: the t statistic is the slope measured in units of its own error. A t of 8.62 means the slope sits more than eight standard errors away from zero — an enormous signal-to-noise ratio, exactly what the high suggested. The two statistics agree.

This 8.62 appears in the regression output in the third column — the coefficients, standard error, and T test columns — corresponding to the row for , because we are talking about . The whole focus of the computation is this single number.

9.3.5 The P-Value Approach and the Decision

To judge the test statistic without a T distribution table, use the p-value approach from the earlier normal distribution discussion: if the p-value , reject the null hypothesis; if , fail to reject the null (accepting it in the loose sense). For the calculated , the p-value is

What a p-value means here. The p-value is the probability of seeing a t statistic at least as extreme as the one observed (8.62) when the null hypothesis is true — i.e., the probability that random chance produces a slope this far from zero when there is actually no relationship. Small p-value means "either the null is false, or a very unlikely coincidence happened"; with , the coincidence explanation is a quarter of a percent — too unlikely to believe, so the null is rejected.

At a significance level of , the p-value is less than , so the null hypothesis is rejected without hesitation. The conclusion: even with a high value, the T test gives the same verdict — there is a relationship between the student population and the quarterly sales. Rejecting the null means accepting that the independent variable contributes — the contribution of exists. So the T distribution approach lets you comment: "there exists a relationship," at least in the context of simple linear regression.

Note on the printed p-value: the professor quotes from the output; the course textbook's software output for the same data prints (rounds below 0.001). The two reports differ only because of software precision — with and 8 degrees of freedom the exact two-tailed p-value is far below 0.001. What matters is the decision rule, and both numbers are far below : reject . On the exam, read and quote the p-value the given output prints; never convert a printed value into something else.

9.3.6 The F Test: Overall Significance

The second approach is the F test, which can be run at the 5% level or the 1% level — either level gives the same result here. The F test uses the ratio of the two mean squares:

where:

  • — the mean square due to regression: the regression sum of squares divided by its degree of freedom, in simple regression. It is an estimate of assuming the relationship exists.
  • — the mean square due to error: , an estimate of that is valid whether or not the relationship exists.

These values come straight from the one-way analysis of variance table you already know: the degrees of freedom column, the sum of squares column, and the mean sum of squares column. For the PJA data, the professor states and , giving

Worked example: the F test arithmetic.

with 1 degree of freedom in the numerator and in the denominator. The F table shows that already leaves area 0.01 in the upper tail with these degrees of freedom; 74.25 is far beyond, so the p-value is far below 0.01.

Sense-check: the F statistic compares two independent estimates of the error variance : the regression-based estimate (inflated if a real relationship exists) and the pure-error estimate. An F near 1 means "the line explains nothing extra"; an F of 74.25 means the line explains far more than pure noise — same verdict as the T test. In fact, for simple regression the two statistics are directly linked: , which matches up to rounding. Both tests always agree in simple regression.

There is no need to consult the F distribution table: the same conclusion is reached by mere inspection of the p-value, which is much, much less than any usual level — the professor reports it as .

9.3.7 Both Tests Unanimously Conclude

Recap: simple linear regression testing comes in two equivalent flavours. The T test looks at the single slope ; the F test looks at the whole regression ; in simple regression they are the same test (), and both reject here with a p-value far below 0.01. Bridge to what comes next: with two independent variables (Section 9.4) the T test and F test stop being interchangeable — the F test checks the overall model while the T tests check each variable separately, and that split is what makes multicollinearity (Section 9.6) detectable.

At this juncture, both the T test and the F test unanimously conclude that there exists a relationship — the linear relationship — guided by the (and multiple ) values. This is exactly why significance testing is required: simply estimating the regression and looking at the value alone does not suffice. The same conclusion stated in the language of the example: a relationship exists between the size of the student population and the quarterly sales.

Real-world connection: this is the standard scrutiny applied to every fitted model in industry — a retail chain checking whether store demographics really predict revenue, a bank checking whether a credit-scoring line has a genuine coefficient before deploying it. In every case the ritual is identical: estimate, read , then prove the relationship with the T and F tests before the model is allowed to influence decisions.

Exam note: a natural question here is the "significance testing of the regression" one — be ready to pull , , , , and the p-values from a regression output and to state the T and F conclusions in words. Expect the professor's pattern of mixing a computation with a comment-on-the-output part. Common exam variant: given , , and the ANOVA table, recompute , , , , and , then state both conclusions.

9.4 Multiple Linear Regression and the Multiple Coefficient of Determination

9.4.1 From One Independent Variable to Several: The Butler Trucking Example

Hook: A trucking manager knows that travel time grows with miles driven — but deliveries take time too. One variable explained 66% of travel time. Could a second variable explain the other third? That question launches multiple regression.

The next step is what happens when more than one independent variable is included, instead of the single variable of simple linear regression. This is a simple extension of ordinary linear regression, and it is introduced through the Butler trucking example (the page reference for the problem is in the textbook).

In that problem:

  • — the travel time (hours) of a truck making deliveries,
  • — the miles traveled,
  • — the number of deliveries made.

If you consider only and the dependent variable , everything you already know from simple linear regression applies unchanged. But the manager felt that the number of deliveries should also contribute to the total travel time. Once you go beyond one independent variable — and you can even work with hundreds of variables — the analysis proceeds differently. The lecture focuses on one dependent variable and two independent variables:

where:

  • — predicted travel time;
  • — miles traveled, — number of deliveries;
  • — the intercept: predicted travel time when both and (the base level before any miles or deliveries);
  • — the effect on predicted travel time of one extra mile, holding the number of deliveries fixed;
  • — the effect of one extra delivery, holding miles fixed.

The phrase holding the other variable fixed is the new ingredient. In simple regression the slope is unconditional; in multiple regression each coefficient is a partial effect — the contribution of that variable after the other variables have already been accounted for. This is the single most important reading skill in multiple regression: is not "what miles do to travel time" but "what miles do to travel time given that we already know the delivery count".

The output side by side: one output is the simple linear regression of on ; the other is the multiple linear regression output of on , using the shorthand notation "X1 plus X2" for the two independent variables in the model.

9.4.2 R-Square Grows: From 0.66 to 0.9037

For the Butler trucking problem, the simple linear regression on alone explains about 66% of the total variation — . When (the number of deliveries) is added, the climbs to . So one observation is immediate: adding an independent variable increased .

The general rule the professor states: adding more independent variables generally increases — it will never decrease. At best it stays the same, otherwise it goes up.

The multiple coefficient of determination is still the same ratio:

only now computed from the ANOVA table of the multiple regression — "where this is SSR and this is SST." With this, the amount of total variation explained by the recommended multiple linear regression model is visible, and the multiple model gives a better idea than the simple one. Note the notation contrast again: small denotes the coefficient of determination for simple linear regression, and capital is used for multiple linear regression.

Why can never fall when a variable is added. Recall with fixed for a given data set. Adding a variable can only ever reduce the error sum of squares: the least-squares machinery can always set the new coefficient to zero and reproduce the old fit exactly, and if a slightly different value fits the data better, shrinks further. Since , a smaller means a larger or equal , so a larger or equal . In Butler trucking, grew enough to take the explained fraction from 0.66 to 0.9037 — the deliveries genuinely add predictive power (the textbook prints the value as 0.9038; the difference from 0.9037 is rounding by the two software packages).

Worked example: the Butler trucking numbers. The multiple-regression ANOVA table reports and (rounded), so

while the simple model on miles alone had . Sense-check: sits between 0 and 1 as it must, and it is larger than the simple as the "never decreases" rule demands. Verbally: adding the delivery count raised the explained variation in travel time from about two-thirds to about 90% — the second variable earns its place in this model.

9.4.3 The Trap: High R-Square Is Not a License to Add Variables

Pitfall — the variable-adding trap: because never decreases, a high is not a license to keep piling variables into the model. The professor's warning is explicit: adding variables may complicate the model, or give you unjustifiably high values. Even a completely redundant variable — one unrelated to the problem — will not lower ; it stays the same or creeps up slightly, while the model only gets more complicated without any contribution from the new variable. Looking at alone, you cannot tell: it may be the same. This is the precise reason a second measure, the adjusted , is recommended — that is the subject of the next section.

The challenge this creates: because of the appealing increase in , you cannot simply keep adding more and more variables to the model. Adding variables may complicate the model, or give you unjustifiably high values. It is not recommended. Even if redundant or unnecessary independent variables are added, you may see no change in — it stays the same or increases slightly — while the model only gets more complicated without any contribution from the new variable. Looking at alone, you cannot tell: it may be the same. This is the precise reason a second measure, the adjusted , is recommended — that is the subject of the next section.

The mental model: is a reward that pays out for every variable added, deserved or not; a model with twenty useless variables will score higher than a clean model with two good ones, even though the clean model is the better science. Model selection so needs a measure that charges rent for each variable — the adjusted of Section 9.5.

9.4.4 Reading the Regression Output Columns

The regression output is organized in columns — coefficients, standard error, T test, and p-value — and the same reading applies in the multiple setting. The T test column corresponds to each independent variable: the row for shows the test for its coefficient. The p-values tell you directly whether to reject the null for that coefficient at your chosen level.

Reading a multiple-regression output row by row:

  • Coefficients column — the estimated : the row labeled Constant holds ; each variable row holds its own slope.
  • Standard error column: the sampling wobble of each coefficient.
  • T test column for each coefficient, testing for that variable individually.
  • p-value column — the probability attached to each t; compare with and reject when .
  • 95% confidence interval columns — for each coefficient, the interval ; the interval and the t test always tell the same story (see the student question in Section 9.4.6).

9.4.5 Running Regression in Excel: The Analysis ToolPak

The questions about how to run these regressions led to a short demonstration with the Excel Analysis ToolPak. This is the everyday workflow practitioners use when they need a regression without dedicated statistics software.

Purpose: produce the standard regression output (coefficients, standard errors, t statistics, p-values, ANOVA table, ) from data sitting in spreadsheet columns.

Inputs and outputs: you feed in the dependent variable column (the range) and the independent variable column(s) (the input range); Excel writes out the same summary-output table read throughout these sessions.

Steps:

  1. Arrange the data. Put the dependent variable in one column (for example ) and the independent variables in the next columns ( and ). Label the columns; the Labels checkbox can then pick up the headers.
  2. Open Data Analysis. Go to the Data tab, where there is an option called Data Analysis. If that option is not visible, invoke it through the Add-Ins button — you enable the Analysis ToolPak add-in, and the Data Analysis button appears.
  3. Choose Regression. Inside Data Analysis, the tool pack offers ANOVA, correlation, covariance, histograms, and so on. Choose Regression and press OK.
  4. Set the ranges. In the Regression dialog, set the Y range to the dependent variable column and the input X range to the independent variable column — only for simple linear regression, both and together for multiple linear regression.
  5. Check Labels and press OK. The Labels checkbox lets Excel pick up the header labels you provided. Press OK and you get the output used throughout these sessions — the same output shown in class.

Between the simple run and the multiple run, the , the coefficients, and everything else differ. That is the entire difference: simple vs. multiple comes down to what you put in the input X range. The same Data Analysis menu also has "Anova: Single Factor" — if you have a one-way analysis of variance problem, you use that option with the same logic: feed the data, run it, read the output. If you are comfortable with Python instead of Excel, the same computation is a simple line of code.

Real-world: this Data Analysis ToolPak workflow is the everyday way practitioners run quick regressions in spreadsheets without any statistical software.

9.4.6 Student Questions and Answers

Q: How do I run the regression in Excel? (Gurudat)

A: Go to the Data tab and open Data Analysis — if it is missing, enable it through the Add-Ins, choosing the Analysis ToolPak. Pick Regression, set the Y range to the dependent variable and the input X range to the independent variable or variables (X1 alone for simple regression, X1 and X2 for multiple), check Labels, and press OK. The output is the same table we read in class; with more variables included, the R square and the coefficients change accordingly.

Q: What does the 95% confidence interval column in the output tell us? (Jyothir)

A: The 95 percent confidence interval for a coefficient is talking about the same thing as the T test. We are validating whether beta1 equals zero or not. If zero is included inside that interval, there is one conclusion; if zero is not included, there is another conclusion — like rejecting or not rejecting the null. So the confidence interval and the hypothesis test agree: zero inside the interval supports beta1 equal to zero, zero outside the interval supports beta1 not equal to zero.

Why they agree: the interval is built as , the estimate plus or minus a margin of error. The t test rejects exactly when is far from zero — which happens exactly when the margin is smaller than the estimate itself, i.e., when the interval misses zero. Same arithmetic, two presentations. (For the PJA slope, the 99% interval is , i.e., 3.05 to 6.95 — zero is far outside, matching the t test's rejection.)

Q: Should I upload the entire Excel sheet with all the output? (Daval)

A: No — instead of uploading the entire information, upload whichever is relevant, spending a little time on it. If everything is uploaded, the evaluator has to verify all of it. The software may give you maximum information, but much of it is not required for the basic questions, because the mid-semester is not beyond these basics. Note that the exam division has been instructed to allow usage of Excel (or other software) during the proctored exam to address the relevant problems.

The three questions cover three different confusion points: the procedure (how to run it), the interpretation (what a printed column means), and the workflow for submission (what to hand in). Note the exam-relevant fact embedded in the last answer: software use is permitted during the proctored exam, so the skill being tested is not remembering the numbers — it is reading the output correctly and selectively.

9.5 The Adjusted Coefficient of Determination

9.5.1 The Problem: R-Square Never Decreases

Hook: You add a completely useless variable to your model — a random number, or something with nothing to do with the problem. Plain does not notice: it stays exactly where it was. A measure that cannot notice a useless variable cannot be trusted for model selection. Enter the adjusted .

The adjusted exists because plain has a blind spot. Adding independent variables never decreases — it stays the same or increases. So even when the added variable is completely redundant — unrelated to the problem, carrying no information — refuses to punish the model. A model can accumulate useless variables and still display an unvarying, high . Judging models on plain alone misleads, and this is why the adjusted measure is recommended.

Why does behave this way? The mechanism was shown in Section 9.4.2: least squares can always reproduce the old fit by setting the new variable's coefficient to zero, so never increases, never decreases, and so never decreases. A redundant variable simply earns the model "no penalty and no reward" — plain cannot tell a useful variable from a useless one, because both leave it unchanged or higher.

9.5.2 The Formula and Its Logic

The adjusted adjusts the coefficient of determination for the number of variables in the model. The formula takes the standard form

where:

  • — the number of observations ("N is the number of observations", as the professor described the ingredients),
  • — the number of independent variables in the model ("P is the number of independent variables which we are considering"),
  • — the plain multiple coefficient of determination,
  • — the adjusted value (written in the textbook; the bar is the lecture's notation).

Why the formula does what it does. Rewrite it as . The numerator is the unexplained fraction — the part of the variation the model fails to explain. The multiplier is a correction that grows as grows: with fixed, every extra variable shrinks the denominator , so the whole fraction gets larger, and subtracting a larger number from 1 lowers .

Two forces are now in competition:

  • Fit: if the new variable genuinely explains variation, falls.
  • Penalty: the added variable grows , which inflates the multiplier.

If the new variable earns its keep, the fit force beats the penalty and rises. If it is redundant, the fit force is absent — does not move — and only the penalty acts, so falls while plain stays put. That is the entire logic: plain sees only fit, adjusted sees fit minus rent for the variables used.

The penalty comes from the degrees of freedom: every extra independent variable reduces , and that lowers the adjusted value unless the added variable earns its place by actually explaining variation.

9.5.3 The Redundant Variable Demonstration

The professor works the logic with hypothetical numbers. Suppose a model with two independent variables, written as , has

Now add a third independent variable — in the Butler trucking context, is the miles traveled and the number of deliveries, and is something unusual with no connection to the problem at all, a redundant variable. After adding it:

  • the redundant variable leaves unchanged — it stays the same,
  • the adjusted drops — from about 0.97 down to somewhere around 0.80.

Careful with the arithmetic — the exact size of the drop depends on and . With observations and variables, the stated closes the first step exactly:

Adding the redundant third variable () while stays 0.971:

The drop is real but small in this particular case, because the multiplier grew only from to . The lecture's "down to about 0.80" describes the dramatic version of the effect, which occurs when the penalty is large — small , many redundant variables, or a partially informative variable that fails to pay for its rent. What is invariant across every case, and what the exam tests, is the direction: plain unchanged or slightly up, adjusted down.

That drop is the trigger. It indicates the redundant variable is being taken care of: the adjusted measure punishes the useless addition, while plain stays silent. The recommendation is firm: never conclude with the simple alone — conclude with the adjusted . With the adjusted value you can say which model is necessary: if a variable contributes nothing, its adjusted penalty shows it, and you can drop that variable from the analysis.

Worked example: a redundant variable added to the PJA model (closing numbers). Start with the simple PJA regression: , , . Its adjusted value:

Now add a completely redundant second variable (say, the street number of each restaurant) so that stays exactly 0.9027, but becomes 2:

Step by step: the unexplained fraction never changes (the useless variable explains nothing), but the multiplier grows from to , so the subtracted penalty grows from 0.1095 to 0.1251, and the adjusted value falls from 0.8905 to 0.8749 while prints 0.9027 in both outputs.

Sense-check: the direction is exactly the professor's point — identical , lower . The adjusted measure has registered the useless addition; plain has not. A comparison of the two numbers, side by side, is the red flag that a variable is redundant.

9.5.4 Student Question and Answer

Q: What happens to the adjusted R square when more independent variables are added?

A: Whenever more independent variables are added, the adjusted R square makes an impact — that is exactly where it differs from the plain R square. Adding a redundant variable leaves the R square the same or slightly up, while the adjusted R square drops, and that drop signals that the extra variable is not contributing.

The answer names the contrast cleanly: "that is exactly where it differs from the plain R square." The two measures respond differently to the same addition — ignores the redundant variable, docks it — and comparing the two responses is how you detect the redundancy. Note the professor's phrasing "makes an impact": the adjusted value always reacts to an addition, up when the variable earns its place, down when it does not.

9.5.5 How to Judge Models: Practical Summary

Recap and exam note: plain only rewards, so it can never notice useless variables; the adjusted charges rent per variable and drops when a redundant variable is added while stays put. Conclude with the adjusted value. Expect a question where you are given , the number of observations , the number of independent variables , and asked to compute or interpret the adjusted — especially the case where a redundant variable is added and the adjusted value drops while does not change.

Many approaches exist — one simple approach is the one delivered here, and you can even discuss all possible combinations of variables if you want to go further. The core habit to take away: compare models on adjusted , not raw . This is also a favorite exam angle: explain why alone misleads when variables are added, and use the adjusted to recommend which variables to keep and which to drop.

Real-world connection: this is exactly how a data team prunes a model before it ships — a credit-risk model starts with dozens of candidate predictors; adjusted (in practice, software reports it as "R-sq (adj)") is one of the standard screens that separates the predictors earning their rent from the ones riding along for free. In the textbook's own multiple-regression example, the two-variable Butler model reports with adjusted — note how the adjusted number sits below the plain one, precisely the "rent" effect.

9.6 Multicollinearity

9.6.1 What Multicollinearity Means

Hook: Your regression output says the overall model is highly significant — and yet every individual variable is insignificant. Both statements are printed on the same page. How can that be? The answer is one of the classic surprises of regression: multicollinearity.

The discussion returns to the Butler trucking example to develop the concept of multicollinearity. Here is the travel time, is the miles traveled, and is the number of deliveries — wherever appears, connect it to miles traveled, and to the number of parcel deliveries.

In the regression model

we simply use the terminology that and are two "independent" variables. But the term independent is only a label: the independent variables may or may not be independent in the statistical sense. At some point the two may be related to each other, and that relationship contributes to our predictions. In practice, independent variables in multiple regression problems are usually correlated with each other to some degree — you cannot say they are completely independent. The question is whether the relation that exists between them makes an impact on prediction or estimation. When it does, you have the multicollinearity problem, and the session covers how to detect it and how to mitigate it.

Multicollinearity, then, is the correlation among the independent variables themselves — the predictors being related to each other — and it matters because it changes how the individual coefficients must be read. The term "independent variable" is a role label (it predicts ), not a guarantee of statistical independence, and confusing the two is the root of much confusion here.

9.6.2 The Petrol Liters Thought Experiment

To make the idea concrete, swap the second variable. Instead of = number of deliveries, suppose = number of liters of petrol consumed by the vehicle. Then it is logically clear that and are related: miles traveled is directly proportional to the petrol consumed, and there will be a high correlation between the two variables. There is no confusion about that.

The professor's analogy: miles traveled is directly proportional to liters of petrol consumed — drive twice as far, burn roughly twice the petrol. The two measurements are the same information wearing two different coats. If you already know the petrol, knowing the miles adds almost nothing; if you already know the miles, knowing the petrol adds almost nothing. So carrying both into the model is carrying the same piece of information twice. Where the analogy breaks: the proportionality is only approximate in reality — city traffic, engine condition, and driving style all disturb the exact ratio — so the correlation is extremely high but not perfect.

The point of interest: should you carry both variables into the model? If both are related to each other, there is no point in carrying both — one of them suffices, because the other is not contributing much to assessing or predicting the situation. In the model

the two predictors are known to be highly related, and the tests reflect that.

9.6.3 F Test for Overall Significance, T Test for Individual Contribution

A key division of labor in multiple regression: the F test is about the overall significance of the regression, while the T test is about the individual contributions of the factors — , , and so on.

Formalize — the two tests ask different questions. The F test tests all slopes at once:

A significant F says "as a group, the predictors carry information about ". The t tests then ask, one variable at a time, vs. — "does this variable carry information after the other variables are already in the model?" That last clause is the whole difference. In simple regression the two questions coincide (one variable: overall = individual, and ); in multiple regression they separate, and highly correlated predictors can make them give opposite-looking answers.

Running the F test on this problem leads to the conclusion that a significant relationship exists — the overall regression is significant. But now run the T tests. For vs. , the T test concludes with accepting the null: there is no significant contribution from , the miles traveled. Put in words: with already in the model, is not significantly contributing — the miles traveled are not making an impact on travel time given that the petrol consumption is already known. And symmetrically, for vs. , the result turns out to be accepting the null again: with in the model, knowledge of the liters of petrol consumed adds no sense for this model.

This is the concept of multicollinearity in action: because the two variables are related, each one individually adds nothing once the other is present, and the model should be revisited — should both variables be carried, or should the model be retained differently?

Worked walkthrough: the petrol variant of Butler trucking. Suppose the fitted model is with = liters of petrol. The output:

  • F test for overall significance: p-value far below 0.01 → reject . Conclusion: the predictors as a group genuinely relate to travel time.
  • t test for : p-value large (say 0.4) → fail to reject . Conclusion: once petrol is known, miles add nothing significant.
  • t test for : p-value large (say 0.5) → fail to reject . Conclusion: once miles are known, petrol adds nothing significant.

Reading the page: the same output reports "the model is significant" and "no variable is significant" — not a contradiction, but the signature of two predictors carrying nearly the same information. Sense-check against the analogy: if knowing petrol makes miles redundant and knowing miles makes petrol redundant, both t tests failing makes complete sense — the group of two explains travel time well, but neither twin can claim credit once its sibling is present.

9.6.4 The Difficulty and How to Detect It

The general difficulty, with many variables and so on: when you test the significance of individual parameters with T tests, multicollinearity makes it possible to conclude that none of the individual parameters are significantly different from 0, even while the overall multiple regression equation is significant. An overall significant model with no individually significant predictors is the signature of the problem — and it is caused by the correlations among the independent variables. The problem is avoided when you anticipate the correlation among the independent variables in advance and omit a few variables.

Pitfall — misreading the paradox: when every t test is insignificant but the F test is significant, the temptation is to conclude "the model is useless" or "the software is broken". Neither is right. The model is significant as a group; the individual coefficients simply cannot be separated from each other's shadow. A further warning from the reference treatment: under severe multicollinearity the estimated coefficients can even take the wrong sign — a predictor known to raise may print with a negative coefficient. Little faith can be placed in individual coefficients when multicollinearity is high.

Several tests exist for determining whether multicollinearity is high enough to cause trouble. The one given here is the correlation matrix: compute the correlation matrix among all the independent variables and inspect it. The rule of thumb:

that is, correlations among , or , or above 0.7 signal that a multicollinearity problem exists in the model.

Visual intuition — reading a correlation matrix. Picture a small table: rows and columns both labeled with the predictor names , each cell holding the sample correlation between the row variable and the column variable. The diagonal always holds 1 (every variable correlates perfectly with itself); the matrix is symmetric (the correlation of with equals that of with ). You scan only the off-diagonal entries — every pair — and apply the rule: any entry beyond in absolute value is a warning flag. In the original Butler trucking data the correlation between miles and deliveries is only — comfortably below the flag, which is why both coefficients stay interpretable there. One-sentence takeaway: the matrix is the map of the problem — a cluster of high off-diagonal values reveals which predictors are twins.

9.6.5 What to Do About It

The remedy at this basic level is a suggestion, to be weighed against other perspectives: every attempt should be made to avoid including independent variables that are highly correlated. Once you anticipate the correlation, drop one of the pair and rerun the analysis with the remaining variables — exactly as in the earlier example, where instead of taking both and , you take any one of them, since both being related means the second cannot make an impact on the model.

Real-world: this is the classic fleet-management situation — miles driven, fuel consumed, and delivery counts are all correlated, and a naive regression will report an overall significant model while every individual coefficient looks insignificant. The practical habit: before trusting a single coefficient, run the correlation matrix on the predictors; when any pair exceeds the 0.7 mark, decide which member of the pair to keep, drop the other, and refit.

Recap and exam note: multicollinearity is correlation among the independent variables; its signature is an overall-significant model (F test) with individually insignificant coefficients (t tests), because each predictor adds nothing once its correlated sibling is in the model. Detect it with the correlation matrix; the rule of thumb flags any pair with sample correlation above 0.7; the remedy is to drop one of the pair. Be ready to explain the F-test-overall vs. T-test-individual distinction, to state what a correlation matrix is used for, and to apply the 0.7 rule of thumb to a given correlation matrix.

9.7 Categorical Independent Variables and Dummy Variables

9.7.1 The Repair Times Problem

Hook: Every variable so far has been a number you could measure — population, miles, deliveries. What do you do when the predictor is a kind of thing rather than a quantity: electrical repair or mechanical repair? Regression only understands numbers, so the question is how to teach it a category.

The next discussion previews a variation on the regression theme, using a problem whose purpose is stated up front: predict the repair times. In this data set:

  • — the repair time, measured in hours,
  • — the months since last service, abbreviated MSLS,
  • — the type of repair.

There are 10 service stations in the data. For each station, the months since last service are recorded, the type of repair is recorded, and the repair time in hours is recorded. The only variation from everything seen so far is the nature of the data for : the type of repair is a categorical variable — it is qualitative, not quantitative. Until now, all the independent variables have been quantitative, continuous variables, and the mathematics handled them without trouble. The problem arises exactly at this point: how do you numerically quantify a categorical variable so it can enter a regression?

A categorical variable (a variable whose values are categories or labels rather than measurements) cannot be added or multiplied, so it cannot sit inside an equation like . The type of repair takes only two labels — electrical or mechanical — and the regression machinery needs one number per observation in each column. The bridge between labels and numbers is the dummy variable.

9.7.2 Introducing the Dummy Variable

The way to handle it is to rewrite the model by introducing dummy variables — binary indicator variables coded 0 and 1. For the type of repair:

In the data, electrical repairs are replaced with 1 and mechanical repairs with 0, producing a column of 1s and 0s (the pattern used in the example runs 1, 0, 1, 1, 0, 0, 1, 1). With this numeric coding, the regression can be run happily.

Formalize — why 0/1 coding works. The dummy is an indicator: it flags which category each observation belongs to. Its coefficient then becomes a switch. In the model

plug in for mechanical repairs:

and for electrical repairs:

So the single equation silently contains two parallel lines: the mechanical line with intercept , and the electrical line with intercept . The two lines share the same slope — the months effect is assumed identical for both repair types — and differ only in their intercept. The dummy coefficient measures the vertical gap between the two lines: the average difference in repair time between an electrical and a mechanical repair, at any given number of months since last service. That is the entire meaning of "a categorical variable enters regression through a dummy".

But the moment categorical independent variables enter the model, a bit of careful interpretation is required — you cannot interpret the results the same way you did for simple and multiple linear regression with quantitative variables. A careful way of carrying out the analysis, and careful interpretation, are needed because of the involvement of dummy variables.

9.7.3 The Two Estimated Models

First, running repair time on alone — a simple linear regression on months since last service — gives the model

Then, including the dummy variable for the type of repair, the estimated multiple regression is

where MSLS stands for the months since last service and stands for the type of repair.

Worked example: reading the two fitted models.

Model 1 — months only. : a station whose last service was months ago is predicted to need hours of repair, with no distinction between repair types. Each extra month since service adds 0.34 hours (about 20 minutes) to the predicted repair time. This model explains only about 53% of the variability in repair time.

Model 2 — months plus type. . Split by category:

  • Mechanical (): .
  • Electrical (): .

The two lines are parallel: the months effect is 0.38 hours per month for both types, and the electrical line sits 1.26 hours higher at every value of — electrical repairs take on average about 1.26 hours longer than mechanical repairs, whatever the service history. Adding the type variable raises the explained variability to about 86%.

Worked prediction for a concrete case: a station with 4 months since last service needing an electrical repair: hours. The same station with a mechanical repair: hours. The gap between the two, 1.26 hours, is exactly the dummy coefficient.

Sense-check: the predictions are positive, plausible, and ordered sensibly — electrical work takes longer, older service history takes longer — matching the professor's example data where repair times run from 1.8 to 4.9 hours.

Writing down the estimated regression model, however, is not enough: the model does not yet account, in an interpretable way, for the mechanical vs. electrical distinction.

The professor's caution: for any given value of and any value of , you cannot simply calculate or predict the repair times — much more analysis is required before predictions make sense. The warning has two edges: first, you must know which of the two embedded lines an observation belongs to before a number means anything (the naive reading "just plug in and " works only after the dummy structure above is understood); second, the full treatment of categorical independent variables — including the careful interpretation and the assumptions on the shared slope — is the scheduled next discussion. The full treatment of categorical independent variables is the scheduled next discussion.

9.7.4 How to Interpret Dummy Coefficients (Reading Ahead)

The careful interpretation that dummy variables demand is previewed by the two models themselves. In the simple model , every extra month since the last service adds 0.34 hours to the predicted repair time, with no distinction between repair types. In the multiple model , the coefficient 1.26 on the dummy shifts the predicted repair time up by 1.26 hours when the repair is electrical (compared with mechanical, where ). Both the intercept-level shift and the interplay with the quantitative predictor are what need careful treatment — and that is exactly the discussion that follows the mid-semester.

Exam note: the professor explicitly states this is the next session's topic — you are expected to know that categorical independent variables enter a regression through dummy variables coded 0/1, and that interpretation is more delicate than for quantitative variables. Expect to recognize the two-line structure a single dummy creates: same slope, intercept shifted by the dummy coefficient.

9.7.5 When the Dependent Variable Is Categorical Too

It is not always the independent variables that are categorical. There are situations where the dependent variable itself is categorical: predicting whether a political candidate wins or loses (the candidate's outcome is the dependent variable, impacted by several factors, some quantitative and some categorical), or whether a high school student is admitted or not admitted to a particular college. When the dependent variable takes only zeros and ones, these are binary classification problems — the same family as classification techniques such as logistic regression. That is the journey of variations ahead: categorical independent variables, and then categorical dependent variables.

Real-world connection: repair-time prediction with a type-of-repair dummy is the standard maintenance-scheduling setup — service fleets predict how long a job will take so they can route technicians; the category (electrical vs. mechanical) is nearly always one of the predictors, and it enters exactly through a 0/1 dummy. The same coding trick appears across industries: region (north/south), shift (day/night), channel (online/offline) — any two-category label becomes one dummy column, and any -category label becomes dummy columns.

9.8 Preview: Categorical Dependent Variables and Logistic Regression

9.8.1 Binary Classification Problems

Hook: The dummy-variable trick handled a categorical predictor. But what if the thing you want to predict is itself a label — win or lose, admitted or not? A linear equation is a bad tool for output that can only be yes or no. The next session's topic, logistic regression, is the answer to that mismatch.

The upcoming discussion concerns situations where the dependent variable behaves like a categorical variable — taking only zeros and ones. Two examples were named: predicting whether a political candidate wins or loses, and predicting whether a high school student is admitted to a college (admitted or not). Both are binary classification problems. This is also the family behind the classification techniques you may have heard of, among which logistic regression is one.

Why does a zero-one dependent variable need a new method rather than just a dummy coding? Because the linear model is not built for it. A predicted value from a regression line is a number on the whole real line — 0.4, 7.2, minus 3 — but "admitted or not" wants an answer that behaves like a probability, pinned between 0 and 1. Fitting a straight line to zero-one data also makes the residuals behave badly: near the ends of the line the model inevitably predicts below 0 or above 1, which is nonsense for a probability. Logistic regression exists to produce predictions that stay inside the valid range. That is the shape of the argument the next session will develop.

Real-world: the canonical example is spam detection. Given an email, classify whether it is spam or not — a binary classification problem. If you work in any of these areas, you can relate your work to the practical settings that will be discussed.

The common thread across the examples — spam or not, win or lose, admitted or not — is that the question is a decision: the outcome has exactly two possible states, and the model's job is to attach a probability to each state. Classification is everywhere in industry practice: fraud detection (fraudulent or not), medical screening (disease present or not), churn prediction (customer leaves or stays), quality control (defective or not).

9.8.2 What the Next Session Will Cover

The plan after the mid-semester: what logistic regression is, and in what way it is helpful; when to use logistic regression; what type of violations in the mathematical and theoretical aspects appear once logistic regression is applied; and the basic difference between linear regression and logistic regression — how they differ and how to handle that.

Preview of the coming contrast — linear vs. logistic: linear regression predicts a continuous number and assumes the relationship has a constant effect at every level of the predictors; logistic regression predicts a probability between 0 and 1, uses a nonlinear (S-shaped) curve to make the prediction, and interprets coefficients as changing the odds of the outcome rather than the outcome itself. The "violations in the mathematical aspects" the professor mentions are exactly the failures of the linear assumptions (constant variance, normality of errors, predictions beyond the valid range) that make the linear machinery wrong for zero-one data.

The discussion of the theoretical aspects of linear and logistic regressions is announced as the picture for the next session, and you are encouraged to come ready: if you are already aware of these ideas, that is well and good; if not, the session will be a fruitful discussion.

The professor's invitation to arrive with prior familiarity is a signal about how the next session will run: it will be a discussion of concepts and contrasts — what each model is for, when each applies, and how they differ — rather than a purely computational session. A useful preparation is to revisit this lecture's regression machinery (estimated equation, significance testing, -style fit measures) because the next session will build the contrast on exactly that foundation: what stays the same, what breaks, and what replaces it.

Exam Guidance Summary

The one-paragraph version: the mid-semester paper has 6–7 questions in two hours, it is open book, and it mixes hand computations with "read the output and comment" questions. Chapters 7 and 8 (named sections only) are the entire scope. Master the Z distribution, chi-square, the T distribution, the ANOVA table, and regression — and practice reading normal distribution tables. One question, about half to one, will sit outside the routine and separate the grades.

  • Scope: the mid-semester covers Chapters 7 and 8 only, and within them only the specified sections. The population proportions sections are omitted; nothing will be asked on proportions or on factorial experiments. No further effort beyond the listed chapters is needed for the mid.
  • Normal distribution: expect a few problems related to the normal distribution; revisit the normal distribution and practice referring to the normal distribution tables before the exam. This is a skill, not just a topic — hypothesis-testing questions will need quick table lookups.
  • Paper format: around 6–7 questions in a 2-hour, open-book paper. The paper mixes questions you must compute with output-related observation questions — you read an output and comment on it. Practice both modes: fluent arithmetic for the computations, fluent output reading for the comments.
  • Grade differentiation: expect roughly half or one question outside the routine, designed to differentiate grades. Answer everything asked; it is not always routine calculations.
  • Master these: the Z distribution, the chi-square distribution, the T distribution, the ANOVA table, and regression. The model paper (from earlier years of the same course) reflects exactly these.
  • Model paper: the shared model paper is a replica of the actual mid-semester paper — no drastic deviations. Use it to see how to address and attempt the questions; rehearse under exam conditions, timed and on paper.
  • Difficulty: the paper is not as difficult as you may anticipate; the expectation is that everyone can handle it.
  • Software: the exam division has been instructed to allow the usage of Excel or other software during the proctored exam. When uploading Excel output, upload only the relevant portion, not the entire sheet.
  • Textbook editions: the 12th and 13th editions have similar sections; only the problems may vary. Either edition works.
  • Regression topics to expect: writing the estimated regression equation from the output; the coefficient of determination and the multiple coefficient of determination; the T test and F test for significance with p-value conclusions; why alone is not enough; the adjusted and the redundant-variable scenario; the F-test-overall vs. T-test-individual distinction and the multicollinearity correlation-matrix 0.7 rule; dummy variables for categorical independent variables (introduced; full treatment after the mid-semester).

Key Industry Applications

  • Real-world: retail sales forecasting — quarterly sales predicted from the student population size (the PJA problem), where 90% of sales variability is explained by the single population driver. This is the standard site-selection analysis a retail chain runs before opening near a campus or mall.
  • Real-world: logistics — travel time predicted from miles traveled and the number of deliveries (the Butler trucking example); adding a second predictor raises the explained variation from 66% to about 90%. Delivery networks use exactly this model to quote route times and plan fleets.
  • Real-world: fleet management — miles driven, fuel (liters of petrol) consumed, and delivery counts are naturally correlated, the textbook multicollinearity setting; correlation matrices and the 0.7 rule are used to catch it before trusting individual coefficients.
  • Real-world: maintenance and service scheduling — repair times predicted from months since last service and the type of repair (electrical vs. mechanical), a categorical predictor handled with a dummy variable. Service organizations use such models to route technicians and set service-level promises.
  • Real-world: spam detection — classifying an email as spam or not, a binary classification problem that motivates logistic regression. The same zero-one pattern appears in fraud detection, medical screening, and churn prediction.
  • Real-world: admissions and elections — college admission (admitted or not) and election outcomes (win or lose) are binary dependent-variable problems of the same family.
  • Real-world: spreadsheet analytics — the Excel Data Analysis ToolPak (with the Analysis ToolPak add-in) is the standard quick way to run simple and multiple regressions and one-way ANOVA in practice, with Python as a one-line-code alternative.

ASM Lecture 9 notes · Regression Analysis: Significance Testing, Multiple Regression, and Categorical Predictors

Advanced Statistical Methods· postgraduate· 2026-08-11

Sections Breakdown

1Mid-Semester Syllabus and Question Paper Expectations

The exam contract: Chapters 7 and 8 named sections only, the 6-7 question open-book paper format, the five topics to master, and student Q&A on scope.

2Simple Linear Regression Recap and the Coefficient of Determination

The estimated regression equation from the output, R-squared as SSR/SST, the three variations SST/SSR/SSE, and why estimation alone is not enough.

3Significance Testing of the Regression: T Test and F Test

Hypotheses on the slope, the t statistic with its standard error, the ANOVA table degrees of freedom, the worked PJA computation, the p-value approach, and the F test.

4Multiple Linear Regression and the Multiple Coefficient of Determination

The Butler trucking example, R-squared growing from 0.66 to 0.9037, the variable-adding trap, reading output columns, and the Excel Analysis ToolPak workflow.

5The Adjusted Coefficient of Determination

Why plain R-squared never decreases, the adjusted formula and its degrees-of-freedom logic, the redundant-variable demonstration, and practical model judgment.

6Multicollinearity

Correlation among predictors, the petrol-liters thought experiment, F-overall versus T-individual testing, the correlation-matrix 0.7 rule of thumb, and the remedy.

7Categorical Independent Variables and Dummy Variables

The repair-times problem, 0/1 dummy coding, the two estimated models and parallel lines, interpretation cautions, and categorical dependent variables ahead.

8Preview: Categorical Dependent Variables and Logistic Regression

Binary classification problems and why linear regression fails for zero-one outcomes; what the next session will cover.

9Exam Guidance Summary

The professor's exam strategy: scope, paper format, mastery list, model paper rehearsal, software allowance, and regression topics to expect.

10Key Industry Applications

Regression in retail site selection, logistics, fleet management, maintenance scheduling, spam detection, admissions, and spreadsheet analytics.

Postgraduate students in Advanced Statistical Methods

Exam Revision Notes

Below is the distilled, exam-ready core. Every entry comes from the full explanation above. Use this section for rapid review; return to the main notes when a point needs more context.

Mid-Semester Syllabus and Question Paper Expectations

Must-know: Mid-semester = Chapters 7-8 named sections only; no proportions, no factorial experiments; 6-7 questions, 2 hours, open book, mixed computation and output-comment questions; master Z, chi-square, T, ANOVA table, regression.

⚠️ Top pitfall: Wandering outside the named sections of Chapters 7 and 8, or studying proportions and factorial experiments that are explicitly omitted.

Self-check: Name the five topics to be thorough with for the mid-semester paper.

Connects to: Simple Linear Regression Recap and the Coefficient of Determination, Significance Testing of the Regression: T Test and F Test

Simple Linear Regression Recap and the Coefficient of Determination

Must-know: Estimated equation y-hat = b0 + b1x from the coefficients column; R-squared = SSR/SST; SST = SSR + SSE; PJA: y-hat = 60 + 5x, R-squared = 0.9027; lower-case r for simple, capital R for multiple regression.

⚠️ Top pitfall: Stopping at a high R-squared: a high R-squared alone does not prove the relationship is real; significance testing is still required.

Self-check: For the PJA data, SSR = 14200 and SSE = 1530: what are SST and R-squared?

Connects to: Mid-Semester Syllabus and Question Paper Expectations, Significance Testing of the Regression: T Test and F Test, Multiple Linear Regression and the Multiple Coefficient of Determination

Significance Testing of the Regression: T Test and F Test

Must-know: t = b1/s_b1 with s_b1 = s/sqrt(sum(xi-xbar)^2), s = sqrt(MSE), MSE = SSE/(n-2); F = MSR/MSE; reject when p <= alpha; PJA: t = 8.62 (p = 0.00258), F = 74.25 (p = 0.000), both reject H0 at alpha = 0.01; in simple regression t^2 = F.

⚠️ Top pitfall: Quoting the wrong p-value: the professor's output prints 0.00258 while the textbook software prints 0.000; both are far below alpha = 0.01 so the decision is the same — never re-derive or alter a p-value printed in the output.

Self-check: For PJA data, SSE = 1530 and n = 10: what are MSE, s, and the error degrees of freedom?

Connects to: Simple Linear Regression Recap and the Coefficient of Determination, Multiple Linear Regression and the Multiple Coefficient of Determination, Multicollinearity

Multiple Linear Regression and the Multiple Coefficient of Determination

Must-know: Multiple model y-hat = b0 + b1x1 + b2x2 with each slope a partial effect; R2 = SSR/SST never decreases when variables are added (Butler: 0.66 to 0.9037); lower-case r for simple, capital R for multiple; Excel Analysis ToolPak: Data tab → Data Analysis → Regression, Y range and input X range.

⚠️ Top pitfall: Interpreting b1 as an unconditional effect: in multiple regression each coefficient is the effect holding the other variables fixed; and treating a high R2 as a reason to keep adding variables.

Self-check: Why can R2 never decrease when an independent variable is added to the model?

Connects to: Significance Testing of the Regression: T Test and F Test, The Adjusted Coefficient of Determination, Multicollinearity

The Adjusted Coefficient of Determination

Must-know: Adjusted R2 = 1 - (1 - R2)(n-1)/(n-p-1): every extra variable p shrinks n-p-1 and grows the penalty unless the variable explains variation; redundant variable leaves R2 same but drops adjusted R2; always conclude with adjusted R2.

⚠️ Top pitfall: The exact size of the adjusted-R2 drop depends on n and p (the lecture's 'about 0.80' is the dramatic illustration; with n = 10, p = 2 -> 3 the drop is small); the invariant, exam-tested fact is the direction: R2 unchanged, adjusted R2 down.

Self-check: A model has R2 = 0.971, n = 10, p = 2. What is the adjusted R2?

Connects to: Multiple Linear Regression and the Multiple Coefficient of Determination, Multicollinearity

Multicollinearity

Must-know: F test = overall significance of all slopes; t tests = individual contribution after other variables are in the model; multicollinearity signature: F significant while all t tests insignificant; rule of thumb: sample correlation above 0.7 between any pair of independent variables flags the problem; remedy: drop one of the pair.

⚠️ Top pitfall: Reading 'F significant + all t tests insignificant' as a broken model — it is the multicollinearity signature; under severe multicollinearity coefficients can even flip sign.

Self-check: In the Butler trucking data, r between miles traveled and number of deliveries is 0.16. Does the 0.7 rule flag a problem?

Connects to: Significance Testing of the Regression: T Test and F Test, Multiple Linear Regression and the Multiple Coefficient of Determination, The Adjusted Coefficient of Determination

Categorical Independent Variables and Dummy Variables

Must-know: Dummy variable: x2 = 1 for electrical, 0 for mechanical; models y-hat = 2.14 + 0.34x1 and y-hat = 0.93 + 0.38x1 + 1.26x2; the dummy coefficient 1.26 is the intercept shift (electrical line sits 1.26 hours higher, same slope); categorical variables enter regression through 0/1 dummies and need careful interpretation.

⚠️ Top pitfall: Plugging values into a dummy-variable equation without recognizing the two embedded lines (mechanical: intercept 0.93; electrical: intercept 2.19); the professor's caution: you cannot simply predict repair times without the full categorical-variable treatment.

Self-check: For x2 = 1 (electrical) and x1 = 4 months, what repair time does y-hat = 0.93 + 0.38x1 + 1.26x2 predict?

Connects to: Multiple Linear Regression and the Multiple Coefficient of Determination, Preview: Categorical Dependent Variables and Logistic Regression

Preview: Categorical Dependent Variables and Logistic Regression

Must-know: Zero-one dependent variables define binary classification problems (spam detection, election outcomes, admissions); linear regression is the wrong tool because predictions escape the 0-1 range; logistic regression produces probability-like predictions and is the announced next-session topic.

⚠️ Top pitfall: Expecting a straight regression line to work for a zero-one dependent variable — its predictions fall outside the valid 0-1 probability range.

Self-check: Name three real-world binary classification problems mentioned in this lecture.

Connects to: Categorical Independent Variables and Dummy Variables

Exam Guidance Summary

Must-know: 6-7 questions, 2 hours, open book; Chapters 7-8 named sections only; master Z, chi-square, T, ANOVA table, and regression; expect normal distribution problems and a grade-differentiating non-routine question.

⚠️ Top pitfall: Studying outside the named sections, e.g., proportions or factorial experiments, which are explicitly omitted.

Self-check: What are the four distributions/tables plus one topic to be thorough with for the mid-semester?

Connects to: Mid-Semester Syllabus and Question Paper Expectations

Key Industry Applications

Must-know: Each regression concept maps to a named industry setting: sales forecasting (R2), logistics (multiple R2), fleet management (multicollinearity), maintenance scheduling (dummy variables), spam/admissions/elections (binary classification).

Self-check: Which industry setting illustrates multicollinearity, and which illustrates dummy variables?

Connects to: Simple Linear Regression Recap and the Coefficient of Determination, Multiple Linear Regression and the Multiple Coefficient of Determination, Multicollinearity, Categorical Independent Variables and Dummy Variables, Preview: Categorical Dependent Variables and Logistic Regression

Was this lecture useful?

Loading comments…
🤖

BitsNotes AI Assistant

Subject Notes Assistant

Configure AI Chat

Choose how to access the chatbot
Have your own API key?

Switch to "Bring Your Own Key" tab above for unlimited access with any OpenAI-compatible provider.

🔑 Enter API key above to fetch live models from provider, or enter model name manually.
OpenAI-Compatible API Support

Choose any provider preset (Gemini, DeepSeek, Kimi, GLM, MiniMax, Qwen, OpenAI, Groq, Ollama, etc.) or enter a custom endpoint URL.

Security & Privacy First

Your API key is sent directly from your browser to your specified provider. BitsNotes servers never store or see your key.