Skip to main content
Advanced Statistical Methods

Introduction to Advanced Statistical Methods

Published: 2026-08-11
Level: postgraduate
Audience: Postgraduate students in statistics

This course is a first session in advanced statistical methods, built for a class of about 445 students spread over 16 contact sessions of 32 hours in total. The session opens with four real-world case studies showing how statistics can be misused, then builds the descriptive foundation you need for everything later: the types of data, measures of central tendency, variance, and a preview of the normal distribution. The recurring promise of the course is simple: what is taught is what will be examined, nothing more. Everything in the notes below follows the order of the session itself — the cases first, then the descriptive building blocks, then the preview of what inference will need.

1.1 Statistical Literacy: How Statistics Can Mislead

1.1.1 Case Study 1 — The COVID-19 Bar Chart That Faked a Decline

Hook: A chart can be 100% accurate in every plotted number and still tell a lie. How? The first case of the session shows a government health department doing exactly that — without changing a single data value.

Any scientific information can be communicated in many ways. One way is a research paper: you publish an article and it becomes publicly available. From a mathematics background, another way is a mathematical model — people build differential equations, publish them, and convey the information through equations and symbols. But for a layman, equations are a very difficult way to communicate. That is why, on top of theoretical discussions and mathematical models, it is always recommended to convey the same thing with pictorial or diagrammatic representations: pie charts, bar graphs, and heat maps. Not all people know differential equations or mathematical symbols, so it is better to give them the information in their own way. Real-world: data visualization is one way to stay informed about developments such as the spread of a virus.

The first case is a bar graph posted around May 2020 — about five months after COVID-19 started spreading around the world — by the Georgia Department of Public Health in the United States. The chart aimed to show the top five counties that had the highest COVID-19 cases in the past 15 days, and the number of cases over time. The class was asked to inspect the chart and find the two major problems in it.

To see the chart in your mind: imagine a bar graph with the date running along the horizontal axis (the x-axis), and the number of reported cases running up the vertical axis (the y-axis). Five bars stand on each date, one per county, and the counties were ordered from the most cases to the fewest. The trap is in that ordering — the bars line the counties up by case count, not by time, and the dates themselves are not in sequence. The class was asked to inspect the chart and find the two major problems in it.

Q: The sequence of the counties is not the same on all the days. A: Correct, that is problem number one. There are actually two major problems. First, the x-axis has no label — the labels are not shown at all, so you cannot tell what each bar stands for. Second, the bars are not ordered in a chronological way: at one point you see 28 April and then again 26 April, with other April and May dates interleaved here and there. The intention of this rearrangement was to show that the number of cases was decreasing from that day to this date — to make the whole graph look as if the count were dipping. That is not true. The sequence of counties was decided by the case order, not by time, so the apparent decline is an illusion. The chart was later readjusted so it could not be read that way.

The visual takeaway: when the dates hop backward and forward — 28 April followed by 26 April, with other dates interleaved — the eye reads a smooth descending staircase, but the staircase exists only in the layout, not in the data. Take the same numbers, put the dates in real chronological order, and the "decline" is gone.

The point of this case: without knowing the background — statistical illiteracy — the public gets misled. For every problem you can draw a diagram, and for every problem you can run a statistical test, but this course is about giving those diagrams and tests some sense: how to minimize that error and how to improve statistical literacy. The problem here is with the lack of knowledge — not knowing how to portray data or which diagram fits the situation — and the fix is to understand statistics well enough to represent results accurately and reach the audience in a proper way.

1.1.2 Case Study 2 — The Truncated Axis That Exaggerates

The misuse of statistics is present everywhere, and news outlets are not the exception. The American news channel Fox News has been under scrutiny several times for its charts. In one recent case, a graph claimed that the number of Americans identifying as Christians had collapsed over the last decade. The trick was in the y-axis: it did not start from zero. The axis ran from 58 to 78, so a drop from 77% of Americans identifying as Christians in 2009 to 65% in 2019 looked dramatic — a 12-point fall in the plotted frame. Draw the same two bars from zero, and the difference between the two decades looks minor.

Picture both versions side by side. In the truncated version, the y-axis (percentage of Americans) runs from 58 to 78, so the two bars — 77 and 65 — reach nearly the top and bottom of the frame, and the 12-point gap eats most of the plot height. In the honest version, the y-axis runs from 0 to 80, so the same two bars sit near the top of a tall frame, nearly equal in length, and the gap looks like what it is: a small change, not a collapse.

Q: Why does that Christians chart look so alarming? A: Because the y-axis is truncated — it starts at 58 instead of zero. When the axis runs from 58 to 78, a fall from 77% to 65% fills most of the frame, making the decline look significant. If you draw the graph from zero to 80, you will not find that much significance between the two decades. Truncating an axis is a damaging way to mislead the public, even if every plotted number is correct.

The lesson here is broader than axis scales: the way you choose the statistical test can also give you blunders. It is always our responsibility to take care of all those things — the balancing act between accurate representation and honest framing is what the entire course is about.

1.1.3 Case Study 3 — The Pie Chart That Totals 193%

During the 2020 presidential run, a news network showed a pie chart for the race in which the percentages totaled 193%. A pie chart must always total 100% if you represent parts in percentage terms, or 360 degrees if you represent them in degrees.

Q: The pie chart is not adding up to hundreds. A: Exactly. If you represent the data in percentages, the total must be 100; if you represent it in degrees, the total must be 360. A total of 193% is simply wrong — even famous news channels do this.

To see why this is absurd: a pie is a circle, and a circle is a whole. Each slice claims a share of that whole, so the slices must account for exactly the whole — no more, no less. A pie with 193% of the circle is like a delivery that weighs nearly twice what was ordered: someone is either careless with the measuring or deliberately inflating the portions.

Such errors come from one of two places. Either it is a lack of knowledge — not knowing how to portray the data or which diagram fits the situation — or it is deliberate: purposefully showing misleading information to grab the attention of the audience or to instigate controversy. We cannot do anything about deliberate misuse, but when the misuse is unknowing, we can minimize it with the help of sound statistical procedures, which is the business of this course. The same two-way split — knowingly versus unknowingly misleading the public — is the session's standing classification of statistical misuse.

1.1.4 Case Study 4 — Doctors vs Gun Owners: Arithmetic Without Context

This case shows how perfectly valid arithmetic can still produce a disastrous conclusion. Take two sets of data. From health statistics for the United States: the number of physicians is about 700,000 (7 lakh), and accidental deaths caused by physicians per year is somewhere around 120,000 (1,20,000). From the Federal Bureau of Investigation (FBI): accidental gun deaths per year across all age groups is about 1,500.

Now work the per-person rates. For physicians, divide deaths by people:

So there are about 0.17 accidental deaths per physician each year. (A second figure, 0.197, was also spoken in the session; it does not come from these two counts, since deaths, not 120,000. The rate that closes with the stated numbers is 0.17.)

For gun owners, the rate quoted was 0.00188. Check the arithmetic: divide 1,500 deaths by a number of gun owners and ask which denominator produces 0.00188.

The denominator that makes the number work is about 800,000 gun owners. The count spoken in the session was 80,000, but that smaller number would give — about ten times larger. The honest statement is: the quoted rate of 0.00188 implies roughly 800,000 gun owners in the comparison.

Worked example — comparing the two rates. The session's claim was that doctors are about 9000 times more dangerous than gun owners. Run the honest ratio of the two per-person rates:

Doctors are about 90 times as deadly per person, not 9000. The spoken working "0.00188 × 9000 ≈ 0.17" does not close arithmetically: , which is about 17 — one hundred times the physician rate, not the same as it. Whatever number the comparison is dressed up with, the point of the case is that the frame of the comparison decides the story: pick a different denominator, pick a different multiplier, and the conclusion changes entirely. Sense-check: the two rates differ by a factor of about 90, so any claim that multiplies the smaller rate into the larger one must use a multiplier near 90 — a claim of 9000 cannot be right.

Q: So are doctors really about 9000 times more dangerous than gun owners? A: No. The arithmetic gives a completely disastrous conclusion if you take it at face value — the natural next step, said with heavy irony, would be to alert your friends and ban doctors before it gets out of hand. But guns don't kill people, doctors do — and the fact is not everyone has a gun, while almost everyone has at least one doctor. Nobody should panic or skip medical attention. The point is that you must ask questions and think carefully about the validity of what people are telling you: why are the rates computed this way, what populations are being compared, and does the conclusion survive inspection? With only the surface numbers, statistics can be made to say almost anything.

That is the famous quote the session closed this case with: if you torture the data sufficiently, it will confess to almost anything. The honest ratio of the two stated rates is only about 90, not 9000 — and that gap between claim and check is the whole lesson. Every number in the doctor-gun comparison could be verified by the reader; the disaster comes from the context-free reading, not from the sums.

Exam note: "If you torture the data sufficiently, it will confess to almost anything" is the moral of the whole opening — a one-line memory hook for why statistical literacy, not formula-tweaking, is the core skill. When a claim rests on computed rates, redo the arithmetic yourself and ask what populations and denominators the comparison is hiding.

1.1.5 The Statistician's Responsibility

The session paired two contrary statements: it is easy to lie with statistics, but it is hard to tell the truth without statistics. We witnessed the first half with the doctors claim — the statistics gave you that information even though it is not true. The second half is the rescue: for real problems, statistics is the tool that lets you bring out the truth and bring out accurate information to the public.

Real-world: that is where the intervention of the statistician and the statistical analyst comes in — working together with company authorities, they review the numbers so the reports submitted to the public are reliable and, to some extent, accurate. Simply drawing diagrams and submitting reports is not enough; the statistician's job is to check whether the picture the numbers paint is the true picture. In practice this shows up in every industry that publishes figures — health agencies, banks, newsrooms, market researchers — where the person who owns the analysis is the last line of defense between a chart and the public.

Pitfalls to carry forward:

  • A chart can lie with fully correct numbers: missing axis labels, scrambled date order, and truncated axes all change the story the picture tells.
  • A pie chart must total 100% (or 360 degrees); a total of 193% means the data was wrong, whatever the intention.
  • A computed rate is only as honest as its denominator and its population: 0.17 per physician and 0.00188 per gun owner come from counts that must be checked, and the ratio between rates is the comparison that matters.
  • Never take a headline statistic at face value — redo the arithmetic, ask what was compared, and see whether the conclusion survives.

Recap + bridge: The four cases all show the same muscle — checking how numbers are framed before believing the conclusion. With that guard up, the course now turns to the raw material the framing works on: the types of data. The next topic asks the first question of any analysis: what kind of data are you holding?

1.2 Types of Data

Hook: Before any calculation, before any chart, one question decides everything: what type of data are you holding? The same statistical method applied to the wrong data type is not merely weak — at some points it is a meaningless way of doing the analysis.

1.2.1 Qualitative Data: Nominal and Ordinal

For any statistical analysis, the first question is what type of data you are holding. Broadly, data falls into two families: qualitative and quantitative. Qualitative data is subcategorized into nominal and ordinal.

A nominal variable (nominal means "in name only") is a way of labeling without any numerical quantification — gender, hair color, and ethnicity are the examples given. A nominal scale has no order and no arithmetic meaning; the labels are just names. Gender is nominal: "male" and "female" are names, and neither comes before the other. Hair color is nominal: "black", "brown", "blonde" are categories, and averaging them makes no sense. The only legitimate arithmetic on nominal data is counting — how many observations fall in each category.

An ordinal variable (ordinal means "ordered") is one where the order matters — first, second, third, or the grades A, B, C. Ordinal data ranks things, but the gaps between ranks are not necessarily equal. The grade A beats B and B beats C, so the order is meaningful; but the step from C to B is not the same distance as the step from B to A — you cannot say the gap is equal, because the labels only promise an order, not a spacing.

1.2.2 Quantitative Data: Discrete and Continuous

The other family is quantitative data, which is the focus once numerical justifications enter the picture. Quantitative data is subdivided into discrete and continuous.

A discrete variable (from the same root as "discreet", meaning separate) takes separate, countable values with gaps between them. The number of students in a class — 445, or 26 — is discrete: you cannot have 445.7 students. A continuous variable can take any value inside a range. Height, weight, and time are continuous: between 5 feet and 6 feet, a height of 5.5 feet, 5.55 feet, or 5.555 feet are all possible. The session also mentioned the interval and ratio scales, but left them for later reading — the idea to hold onto is that quantitative measurements sit on either an interval or a ratio scale, the difference being whether a true zero exists (temperature in Celsius is interval — 0 is not "no heat"; height is ratio — 0 height means nothing).

Real-world: nowadays data is information — you may not just expect numbers; you may expect text, images, or what-not. Whatever form it takes, it is mandatory to understand the nature of the data before making any data analysis. Broadly, you will come across nominal, ordinal, discrete, and continuous data types, and each specific data type has its own suitable statistical technique.

Data type Family Order? Arithmetic? Example
Nominal Qualitative No Counting only Gender, hair color, ethnicity
Ordinal Qualitative Yes, but gaps unequal Ranking only Grades A, B, C; first, second, third
Discrete Quantitative Yes Full arithmetic on separate values Number of students in a class
Continuous Quantitative Yes Full arithmetic on any value in a range Height, weight, time

1.2.3 Matching the Data Type to the Technique

The essential rule: for any data you can apply any statistical method, but it need not be reliable — at some points it will be a meaningless way of doing the analysis. The relevant question is always: is this a nominal problem or an ordinal problem? If it is, what is the relevant technique to address it? And once you have the results, you have to infer from them.

Later in the course — especially in bivariate and multivariate techniques such as regression analysis — we will first categorize what type of data is being used and how it is ordered, and only then choose the suitable statistical technique. The next session will bring concrete examples of each data type and how to handle them.

Pitfalls:

  • Treating labels as numbers: averaging a nominal variable (the "average hair color") is a meaningless result, even though the software happily prints it.
  • Treating ranks as equal steps: ordinal data like grades A, B, C tells you the order, not the spacing — the gap between grades is not necessarily equal.
  • Forgetting that the choice of technique follows the data type: a method tuned for continuous measurements can be meaningless when forced onto nominal categories.
  • Confusing discrete with continuous: a count (445 students) can never take values between its steps; a measurement (time, weight) can.

Recap + bridge: Data is either qualitative (nominal, ordinal) or quantitative (discrete, continuous), and the suitable technique is chosen after classifying the data, not before. The next topic takes the first summary step that works on quantitative data: describing where the center of the data sits.

1.3 Measures of Central Tendency

1.3.1 Mean, Median, and Mode

Hook: If you could ask a data set one question, it would be "where do you live?" — where is the center of all these values? Mean, median, and mode are three different answers, and each is right in different situations.

With any specific raw data, your initial analysis asks two questions: how does the central behavior of the data look, and what is the spreadness of the data. Mean, median, and mode answer the first question — they all convey the central behavior, but different situations demand different measures.

The mean is the average — it takes into consideration all the data points spread across the data set and takes the average. The median is the middle value when the data is ordered. The mode is the number that occurs the most — the value that keeps recurring.

Measure What it finds Data needed Sensitive to extreme values?
Mean The average of all values Every value Yes — one huge value pulls it
Median The middle value after ordering Ordered list No — only the middle position
Mode The most frequent value Frequency of values No — it ignores almost everything

The median needs one extra detail that the session left implicit and the textbook spells out: with an odd number of observations the median is the middle value itself; with an even number there is no single middle value, so the convention is to average the two middle values. For 445 students, the 223rd ordered value is the median.

1.3.2 Why the Mean Is the Most Reliable Measure

The session asked which of the three measures you would rank first, and why the mean is the best among them. The answer: the mean uses every single value in the data set, so every individual contributes to the result.

Q: Which of the three measures is the most reliable, and why do we rank the mean first? A: The mean, because it takes into consideration all the data points which are spread across the data set and takes the average. Consider the whole class: if we calculate the evaluations and the average comes out around 71, that average counted each individual to reach 71. The mean is the only measure where every individual's contribution matters.

Formalize — the mean. Let be the number of values in the data set, and let (read "x-sub-i") be the -th value, so the data set is . The sample mean, written (read "x-bar"), is the sum of all values divided by their count:

The sum runs over every value from the first to the last — that index bound is what makes the mean "count every individual". The factor then spreads the total evenly over the observations. No value is skipped, none is weighed more than another.

Worked example — the class average. Take the class strength of 445 students. Once all evaluations are in, the average is calculated around 71. Work it backward to see what the mean really did:

The 445 individual scores added up to about 31,595, and dividing by 445 returns the mean:

The mean is 71. The single number 71 carries the contribution of every one of the 445 individuals — a score of 100 and a score of 40 both move it. Sense-check: the average must lie between the smallest and largest score, and 71 sits inside any plausible evaluation range.

The mode, in contrast, can lie about the center:

Q: Could the mode ever report the class average for us? A: No — the mode is nothing but the number of times a value recurs. If a few students — 10 out of 40 or 50 — got 100 out of 100, the mode would report 100 as the average. That is completely wrong, because the mode does not count each and every individual's contribution. You cannot say the class average is 100 because the top 10 students scored full marks; new people will not appreciate that conclusion either.

So among all the measures of central tendency, the mean gives you the best picture of the data — which is exactly why the theory of estimation builds its inference on means.

Pitfalls:

  • Letting the mode speak for the whole class: the mode only counts recurrences. Ten perfect scores out of fifty make 100 the most frequent value, and the mode would cheerfully report 100 as the "average".
  • Forgetting that the median ignores magnitudes: the median only needs the middle position, so it does not feel a huge outlier — useful in skewed data, but it throws away information the mean keeps.
  • Assuming one measure fits every situation: income data, for example, is usually summarized by the median precisely because a few huge incomes pull the mean away from what most people earn; the session's point is that for class evaluations the mean is the most reliable because every individual is counted.

1.3.3 From Means to Estimation Theory

Because the mean is the most reliable summary of the center, the standard inferential tools are built around it. This is a direct bridge to later parts of the course: in the theory of estimation you come across the Z test and the T test, and the inferences on means come into the picture. The Z test and T test both take a sample mean , compare it with a claimed population value, and ask whether the gap is too large to be chance — so the reliability of the mean is what makes those tests trustworthy.

Exam note: Expect the reasoning of this section — why the mean beats the mode for summarizing a large class — to reappear as conceptual groundwork for those tests. The deeper message: a measure that ignores part of the data (like the mode ignoring everyone except the most frequent score) cannot be the foundation of inference.

That deeper message points straight at the next section: a single number that only says "where is the center" still cannot tell you whether the data hugs that center or scatters far from it.

Recap + bridge: The mean is the most reliable measure of the center because it counts every individual — and that is why estimation theory is built on means. But the center is only the first question of analysis; the second question is how the data spreads around that center, which is the next topic.

1.4 Variance and Dispersion

1.4.1 Why the Mean Is Not Enough

Hook: Two treatments, two different density tables, and both average exactly 70 beats per minute. The mean cannot separate them. Which one do you prescribe? That single question is why variance exists.

The motivating question of this section: why is variance required when the mean is enough? The mean tells you where the center is; the second question of any initial analysis is the spreadness — the variability, or dispersion, of the data. Range and variance are the measures that talk about this variability. A student supplied the intuition:

Q: Do you have any examples where variance really matters? A: Yes — in production, when you are deciding on certain parameters while manufacturing units. When you want to estimate the size, say the size of chairs or tables, you need to know how much the measurements scatter. Another example is blood pressure: variance tells you how accurate your average is, or how close your average is to the entire data. To understand that closeness, you use variance as your parameter.

That is the vocabulary of this section: spreadness, dispersion, and variability all describe how far the data points sit from the center. Variance is the single number that captures it. In manufacturing, the point is easy to feel: a chair that measures 40 cm on average is fine, but not if half the chairs come out at 30 cm and half at 50 cm — the average hides the scatter, and the scatter is what decides whether the product fits its purpose.

1.4.2 Worked Example — Old Drug vs New Drug

Here is the full setup that shows why variance is essential. A drug is used to maintain a steady heart rate in patients who have suffered a mild heart attack. Let denote the number of heartbeats per minute recorded per patient when the old drug is administered. On a similar note, suppose you wish to compare a new drug with the old one; let denote the number of heartbeats per minute when the new drug is administered. So and measure the same thing — heartbeats per minute — under two different drugs.

Both drugs were summarized by a density table: for each possible heartbeat rate, the probability that a patient lands on that rate. The average heartbeats per minute is then computed with the expected-value formula. The verbal description given was: "the expected value of X is summation X into f(x)."

Formalize — the expected value. For a discrete variable with probability of taking the value , the expected value — the probability-weighted average — is

Here is the probability of observing heartbeats per minute, the sum runs over every possible rate in the density table, and (read "E of X") is the mean of the distribution. The same idea as the class average, with probabilities as the weights: each rate contributes its value multiplied by how likely that rate is. In the reference treatment this is the "long-run average": run the experiment many times, and the average of the observed rates settles at . Standard texts write the same object as or — the notation is used here to match the lecture.

The old-drug density table shown in the session starts with the rates 40 and 60 at probabilities 0.01 and 0.04; the middle rows of the table were shown on screen. The completed table below keeps every spoken number and reproduces every stated result:

(old drug) 40 60 69 70 71 80 100
0.01 0.04 0.10 0.70 0.10 0.04 0.01

Worked example — the average for the old drug. Multiply each rate by its probability and add:

The average heartbeats per minute for the old drug is 70. Sense-check: the probabilities add to , so the table is a valid density table, and the weighted average 70 sits between the smallest and largest rates, exactly where the center of the table lies.

The class was then asked to compute the average for the new drug in a similar fashion. The spoken working was "40 into 0.40, plus 16 into 0.05, and this is 100 into 0.04 — sorry, 0.40" — the "16 into 0.05" is a slip for 60 into 0.05, and the last term was corrected from 0.04 to 0.40:

(new drug) 40 60 69 70 71 80 100
0.40 0.05 0.03 0.04 0.03 0.05 0.40

Worked example — the average for the new drug.

The new drug also averages 70 heartbeats per minute. Sense-check: the probabilities again add to , and the weighted average is 70 — but notice where the probability sits: the new drug piles 0.80 of its probability at the extreme rates 40 and 100, while the old drug piles 0.70 of its probability at the middle rate 70.

Now the dilemma:

Q: Both drugs give the same average heartbeats per minute — so which one is better? A: That is exactly the problem: with averages alone we are not in a position to decide whether to keep the old drug or switch to the new one. The average does not give us much information here because both are the same. If you look closely into those density tables, the answer appears: the variance is what we should look at next.

1.4.3 The Tolerance Argument and the Verdict

Physicians can tolerate, for a healthy individual, a heart rate within plus or minus two of the average heartbeats per minute — so around 70, the acceptable range is 68 to 72. Now inspect the two density tables against that band. With the old drug, 90% of the people fall within the average heartbeats per minute range (the probabilities at 69, 70, and 71 add to ), and only 2% of the people react extremely — the average of 70 coming down to 40 or shooting up to 100. With the new drug, the picture flips: about 10% of the people sit within the acceptable range, while 80% of the people react extremely, landing at the extreme rates.

How can you capture this information with a single number? Variance.

Formalize — the variance. The variance measures each rate's distance from the mean, squares that distance so that gaps above and below the mean do not cancel, and weights the squared distances by probability. Let be the mean of the distribution:

Every symbol: (mu) is the mean of the distribution, 70 for both drugs; is the deviation of a rate from the mean — how far that rate sits from the center; squares the deviation, so a rate 30 below the mean and a rate 30 above it both count the same distance 900; and weights each squared deviation by how likely the rate is. This is the standard definition given in the reference treatment (the variance is a weighted average of the squared deviations). Some texts use the equivalent form , which follows from expanding the square; both give the same number. The positive square root of the variance is the standard deviation, .

Worked example — computing both variances. For the old drug, every rate deviates from 70: 40 deviates by , 60 by , 69 by , 70 by 0, 71 by , 80 by , 100 by . Squaring and weighting:

For the new drug, the same deviations are weighted by its table:

The comparison settles it: . Sense-check: the new drug's variance is about 28 times larger because its probability sits at the far rates 40 and 100, whose squared deviations are 900 each — exactly the "extreme reactions" the doctors want to avoid.

The old drug's heart rates cluster tightly around 70; the new drug's rates scatter out to the extremes. The lesser the variance, the better the drug — it is more efficient in the sense that it is not widely spread, and it does not give you extreme reactions. With the numerical values of the two variances, you recommend the old drug: not by simply writing it down, but justified with numbers. This is also the pattern you will meet whenever software output is on the exam table: the program gives you the averages and the measures of spread, and your job is the inference — which number is smaller, what that means, and which choice it supports.

Two closing remarks from the session. First, in any distribution — including the normal distributions that follow — these two parameters, mean and variance, are the ones you talk about; higher-order parameters exist but are out of scope for this course. Second, the intuition to remember: a low variance means the data hugs the average, a high variance means it scatters — and when two summaries have identical averages, the spread is the deciding factor.

Scope and assumptions:

  • The expected value and variance formulas used here are the population versions: the density table is the probability distribution of the whole population of patients, so is the true mean, not an estimate from a sample. When you work from a sample instead, the mean and variance are estimated from the data — a distinction the course picks up in later sessions.
  • Variance measures spread around the mean only. If the mean is not a sensible center for the data (for example, a lopsided distribution with a long tail), a single variance number still exists but tells you less.
  • Because deviations are squared, extreme values dominate the variance: one rare extreme rate can make a variance much larger than the rest of the table suggests. That is a feature here — the extreme reactions are exactly what the doctor fears — but in other settings it can hide the typical behavior.

1.4.4 Student Questions and Answers

Q: In the drug discussion, we had 90% on the old drug and 80% on the new drug — so from that, we have to take the minimum value, right? A: No — you calculate the variance with reference to the average. Looking at the density table, somebody can interpret it by eye, but the variance is what gives you a single number with that clarity: it measures how far the heart rates spread around the threshold average of 70. Compute the variance of X and the variance of Y, see that X gives the lesser variance, and then you are convinced that the old drug is the recommendation because it is not giving you extreme reactions. That is how you justify the decision with numerical values.

The exchange repeats the core point: you do not compare the 90% and 80% figures directly — you feed the same density tables into the variance formula, get one number for each drug, and let the comparison of those two numbers decide. The 90% and 80% figures were already part of the reasoning that built the tables; the variance condenses all of that into the two numbers 26.2 and 730.06, and the smaller one wins.

Recap + bridge: When two summaries share the same average, the spread is the deciding factor — variance is the single number that captures spread with reference to the mean, and the lesser variance picks the more efficient drug. The session now steps back to preview the most famous distribution in statistics, whose entire shape is described by exactly these two parameters — mean and variance: the normal distribution.

1.5 The Normal Distribution Preview

1.5.1 Gauss and the Gaussian Name

Hook: Why do heights, test scores, and measurement errors all pile up in the middle and thin out at the ends? The answer is the most fascinating and widely used distribution in all of statistics — and it carries the name of one man.

The session paused before inferential statistics to revisit the normal distribution, calling it a very essential part of understanding estimations, hypothesis testing, and p-values. It is one of the most fascinating and widely used distributions among the discrete and continuous distributions. The distribution is associated with Carl Friedrich Gauss (1777–1855), the German mathematician who worked with it extensively in astronomy and geodesy — fitting curves to noisy measurements — which is why the name stuck. Real-world: in many industries, people even call the normal distribution the Gaussian distribution or a Gaussian. In engineering, finance, and machine learning, you will hear "assume a Gaussian" far more often than "assume a normal distribution" — the two names mean the same bell-shaped curve.

The session shared Gauss's own framing with a joke: God loves average looking, that is why he makes so many of them. Is it true? Yes. The two extremes are always rare while middlings are always common — that is the entire idea.

1.5.2 The Bell Curve: Extremes Are Rare, Middles Are Common

The normal distribution is a bell-shaped curve. The intuition was built with height: if X is height, the tallest person — maybe 16 feet — is very rare; one or two will be there. Two feet or three feet is also rare; very few will be there. So the majority of people lie in between a certain possible range.

To see the curve: draw the horizontal axis (the x-axis) for the value of the variable — height in feet — and the vertical axis (the y-axis) for how many people have that height. The curve rises smoothly from near zero at the far left, climbs to its highest point at the middle, and falls back to near zero on the far right. The peak sits at the average: most people cluster at a middling height. The two tails are the landmarks: the far left and far right are nearly flat because a 16-foot person and a 2-foot person are both nearly impossible. The one-sentence takeaway: the curve is tallest where the average is, and it dies out toward both extremes — the middlings are common, the extremes are rare.

The same joke applies to grading. The instructor loves average looking too — that is why many B, C, D grades are given while A and E are rare.

Exam note: If everyone performs very well, that is fine, but generally it does not happen for large samples — so you can expect the grade distribution of a large class to roughly follow the bell. The point is not that grades must be forced onto a curve; it is that large sets of naturally varying scores tend to concentrate in the middle.

1.5.3 Large Samples and the Law of Large Numbers

The connection to sample size: this course's registered students number 445, while in the previous semester another course had a strength of just 26. Generally speaking, large samples will always follow a normal distribution; small samples may or may not follow a normal distribution. That is what the law of large numbers is all about.

The law of large numbers says, in plain words: as the number of observations grows, the behavior of the group settles down — averages stabilize near their true value, and the spread of the observations takes on a stable, bell-shaped pattern. With only 26 students, one or two unusual students can skew the whole picture; with 445, no single student can dominate, so the collection behaves like the smooth bell. With 445 students, once the evaluations are complete, we will be in a position to decide whether the class behaves like a normal distribution or a completely skewed distribution — but the number suggests it should tend toward the normal.

1.5.4 Why the Normal Distribution Matters in This Course

The normal distribution is the foundation for the inferential part of the course: estimations, hypothesis testing, and p-values all rest on it. A practical note for the entire course: all software output gives you the result directly — there is no explicit way of computing p-values by hand, because Python, SPSS, SAS, or R will give you a direct p-value. Your role is to interpret that output. Still, the course will show how the p-value is computed, because that understanding is also required.

Study advice: start reading chapters 6, 7, 8, and 9 — the material on the normal distribution, estimations, hypothesis testing, and p-values — before the coming sessions; these will be covered briefly, no need to mug it up. The next session takes 30 to 40 minutes to walk through the normal distribution and then introduces the actual course content.

Pitfalls to keep in view:

  • "Large samples always follow a normal distribution" is a rule of thumb for the course, not a guarantee for every data set — a strongly skewed population can still produce a lopsided sample, which is why the professor says the data will decide.
  • The two parameters that describe a normal distribution are the mean and the variance; the session flags that higher-order parameters exist but are out of scope — do not go hunting for them in the exam.
  • Software prints p-values directly; the trap is skipping the interpretation step, because the exam asks for your inference on the output, not the arithmetic behind it.

Recap + bridge: The normal distribution is the bell-shaped home of the middlings — extremes are rare, the mean and variance describe its whole shape, and large samples tend toward it by the law of large numbers. Everything inferential that follows — estimations, hypothesis testing, p-values — builds on this curve, and the next session walks through it properly.

1.6 Course Tools, Textbooks, and Evaluation

1.6.1 Software Tools: R, Python, SPSS

The course shows analyses in R — the sessions show how R runs and what the R output looks like. Students comfortable with Python may explore the same material in Python; the same results will come, so there is no issue. SPSS and SAS were also named as the industry-standard tools that produce direct p-values. The assignment component has no tool restriction: you can use R, Python, SPSS, or your own way — there is no compulsion.

Tool Role in the course What you get from it
R Demonstrated in the sessions Full analysis output, direct p-values
Python Optional parallel track The same results as R
SPSS Industry standard, named in the course Direct p-values, point-and-click workflow
SAS Industry standard, named in the course Direct p-values

Q: Will we be using R or Python, or is it just Python, or one of them? A: I will go through with R in the sessions. You can explore Python or SPSS in your own way — nothing wrong, no issue; the same results will come. For the assignment, use whichever you are comfortable with.

Q: Do those who do not know R or Python need to take a separate course, or will you cover it here? A: It is covered here. There is no such prerequisite for this course, and you may not expect that as a point. If you already know the tools, that is well and good; if not, you will see how R runs and how to read its output during the sessions. From the exam perspective, what matters is your inference on the output — retrieving the numbers behind the screen is not the point of the assessment.

The second exchange carries a rule that runs through the whole course: the software is a means, not the subject. The exam asks what the output means, not how the software computed it.

1.6.2 Textbooks and Study Resources

Three textbooks are used and will be referred to as T1, T2, and T3. T1 covers hypothesis testing — the fundamental material required before doing any analysis, which the course will revisit. About 60 to 70% of the course content comes from T1. T2 is the multivariate analysis textbook, aligned with the course notes. T3 has an SPSS focus. The edition matters: the reference is the 30th edition. If soft copies are available — on the e-library portal or through students who have access — they will be shared with everyone.

Q: Can you tell us which books are recommended for this course, so that we can refer to them later? A: Yes — I started the session with the case studies and will come to the textbooks after that. Three textbooks will be listed: the first on hypothesis testing, the second on multivariate analysis, and the third with an SPSS focus. Please have a look at them and keep them with you.

Q: Will the textbooks be very important, or are the handouts and the class notes enough to appear for the exams? A: The textbook is required, because I will not type the entire question — I will just give you from which chapter, which page number, and which problem it comes from, so you can do the practice problems. If you go through the sessions and the note materials, that is enough; the textbook makes sense if you want more information. Whatever is delivered in the sessions is what the exams focus on — whichever is not taught will not be asked.

Q: Does the edition of the textbook matter? A: The edition matters — I am referring to this 30th edition. If I have a soft copy of any one of them, I will try to share it, and if anybody can get access, please share the links with all your colleagues.

1.6.3 Evaluation Structure and Doubt Support

The evaluation plan for this course: two quizzes and one assignment, with the portions for quiz one, quiz two, and the assignment informed well in advance, and dates posted on the portal in a timely manner. With about 445 students, individual assignments are not feasible for checking, so the assignment will be done in groups: groups can be formed by the students or decided by the instructor, the group size will be mentioned, and each group performs a statistical analysis and submits a report. The point of the group format is to appreciate teamwork.

There are regular and makeup examination slots — two different slots. Makeups are given only for genuine cases, and for on-campus students the concern is scrutinized before approval. The personal advice was direct: do not go for the makeup exams, there is no point, by that time something else will be piled up; there is no such difference between the regular and makeup tests as such. The sessions number 16 contact sessions of two hours each.

The difficulty of the mathematics was addressed head-on: calculus is not severe — you have done algebra, and only a bit of linear algebra is required somewhere; no full calculus solving is expected.

Q: These books do not cover R. If somebody wants to learn R, can you refer something? A: I will tell you whichever links are required and how to do it — not for everything, only when a technique needs it, for example cluster analysis. Whichever technique is required, I will do that. Once you know the concepts, you will explore the tools with your own domain if it is required.

Q: How will you conduct the doubt-clearing sessions? Some instructors create a Telegram group where we can post questions, or use a forum like Taksila. A: I will share something like that too — I will set up a group. If not, you can write me an email and I will reply to that.

Q: What will be the complexity of calculus in this course? A: Not severe. You have done algebra — we are not going to that extent of full calculus. A bit of linear algebra is required somewhere, that is all; no detailed mathematics solving in this course. We will take care of it.

Recap: Two quizzes and one group assignment make up the evaluation; R runs in the sessions with Python, SPSS, and SAS as honest alternatives; T1 supplies the bulk of the course content; and the exam tests your inference on output, not your arithmetic. With the logistics in place, the course proper begins — the normal distribution is the first real topic of the coming sessions.

Exam Guidance Summary

Everything examinable in this first session is collected here. The standing rule from the opening promise: what is taught is what will be examined, nothing more.

  • Coverage policy (most important): the exams focus only on what is taught in the sessions. "Whichever is not taught, I will not ask." Do not search elsewhere — the delivered material is the boundary of the exam.
  • Textbook requirement: the textbook (T1, T2, T3) is required because exam questions will be quoted as chapter, page number, and problem number rather than typed out in full. Sessions and note materials are enough for the taught content; the textbook adds practice problems.
  • Assessment structure: two quizzes and one assignment. Portions for quiz one, quiz two, and the assignment are announced in advance; dates are posted on the portal in a timely manner.
  • Assignment format: group work — groups formed by students or by the instructor, size to be announced; each group does a statistical analysis and submits a report. No tool restriction (R, Python, SPSS, or your own way); teamwork is the point.
  • Exam answer style: expect to give your inference on software output. R, Python, SPSS, and SAS give direct p-values — your role is to interpret them; the course also shows how the p-value is computed, so that understanding matters too.
  • Regular vs makeup: two different slots; makeup only for genuine cases. The advice given: do not plan for the makeup — there is no difference between the regular and makeup tests, and makeup time will pile up work.
  • Prerequisites: no programming prerequisite for the course; calculus complexity is not severe — a bit of linear algebra is needed somewhere, no full calculus solving.
  • Study advice: start reading chapters 6, 7, 8, and 9 (normal distribution, estimations, hypothesis testing, p-values) before the coming sessions; they will be covered briefly — no need to mug up. The normal distribution is the fundamental background for everything inferential in the course.
  • Conceptual groundwork: the opening case studies (misleading bar charts, truncated axes, pie charts that do not total 100%) carry the course's recurring theme — accurate, honest representation of data — and the mean-versus-mode reasoning explains why inference is built on means (Z test, T test).

Key Industry Applications

The case studies and concepts of this session map directly onto real work in health communication, media, manufacturing, and pharmaceuticals:

  • Health statistics and public communication: the Georgia Department of Public Health's COVID-19 bar chart (May 2020) shows how a government body can mislead with an unlabeled, non-chronological chart — and how the corrected chart restored honest communication.
  • News media scrutiny: the truncated-axis chart about Americans identifying as Christians (77% in 2009 to 65% in 2019, plotted from 58 to 78) and the 193% pie chart of the 2020 presidential run show that even famous news networks publish misleading visualizations, deliberately or unknowingly.
  • Crime statistics: the Federal Bureau of Investigation's (FBI) gun-death statistics were used to build per-person rates — a cautionary example of arithmetic without context leading to an absurd conclusion.
  • Manufacturing and quality: variance matters when manufacturing units — estimating the size of chairs and tables — because you need to know how much measurements scatter around the target.
  • Pharmaceutical trials: the old-drug vs new-drug comparison (heartbeats per minute after a mild heart attack) is a miniature clinical-trial decision: two drugs with identical averages are distinguished by variance, and the lower-variance drug is the efficient, recommended one.
  • Software tools in industry: R, Python, SPSS, and SAS are the tools that compute p-values directly; the industry expectation is that analysts interpret the output, not compute it by hand.
  • The Gaussian in industry: the normal distribution is widely called the Gaussian distribution in industry, and large-sample data (like a 445-student class) is expected to follow it via the law of large numbers.

ASM Lecture 1 notes · Introduction to Advanced Statistical Methods

Advanced Statistical Methods· postgraduate· 2026-08-11

Sections Breakdown

11.1 Statistical Literacy: How Statistics Can Mislead

Four real-world cases where charts and rates mislead with fully correct numbers: a scrambled-axis COVID-19 bar chart, a truncated y-axis, a 193% pie chart, and the doctors-versus-gun-owners rate comparison.

21.2 Types of Data

The two families of data — qualitative (nominal, ordinal) and quantitative (discrete, continuous) — and why the statistical technique must follow the data type.

31.3 Measures of Central Tendency

Mean, median, and mode; why the mean is the most reliable summary of the center because it counts every individual, and why estimation theory is built on means.

41.4 Variance and Dispersion

Variance as the measure of spread around the mean, developed through the old-drug versus new-drug heart-rate example using the expected-value formula.

51.5 The Normal Distribution Preview

The Gaussian bell curve, extremes rare and middlings common, the law of large numbers, and why the normal distribution underpins estimation, hypothesis testing, and p-values.

61.6 Course Tools, Textbooks, and Evaluation

R, Python, SPSS, and SAS as the analysis tools; the three textbooks T1, T2, and T3; and the evaluation structure of two quizzes and one group assignment.

7Exam Guidance Summary

The standing exam rules of the course: what is taught is what is examined, questions quoted from the textbook, and answers given as inference on software output.

8Key Industry Applications

Where the session's cases and concepts appear in health communication, news media, crime statistics, manufacturing, pharmaceutical trials, and industry software.

Postgraduate students in statistics

Exam Revision Notes

Below is the distilled, exam-ready core. Every entry comes from the full explanation above. Use this section for rapid review; return to the main notes when a point needs more context.

Statistical Literacy: How Statistics Can Mislead

Must-know: Statistics can mislead with fully correct numbers: check axis labels, axis origin, totals, and the denominators behind every computed rate; the statistician's job is honest framing.

⚠️ Top pitfall: Believing a computed comparison at face value: the spoken claim of 9000 times does not close arithmetically (0.00188 × 9000 ≈ 17), while the honest ratio of the stated rates is about 90.

Self-check: Why did the Georgia chart fake a decline even though every bar value was correct?

Types of Data

Must-know: Classify data before analysis: nominal (labels, no order), ordinal (order, unequal gaps), discrete (countable), continuous (any value in a range); each type has its own suitable technique.

⚠️ Top pitfall: Applying any statistical method to any data type: the result can be a meaningless analysis even though the computation ran without error.

Self-check: Which data family does gender belong to, and why can you not average it?

Measures of Central Tendency

Must-know: The mean is the most reliable measure of central tendency because it uses every value in the data set; a measure that ignores part of the data cannot be the foundation of inference.

⚠️ Top pitfall: Letting the mode report the center: with 10 of 50 students scoring full marks, the mode would report 100 as the average even though most students scored far lower.

Self-check: Why does the mean beat the mode as a summary of a 445-student class?

Variance and Dispersion

Must-know: Variance measures spread with reference to the mean; when two distributions share the same average, the one with lesser variance is the efficient, recommended choice.

⚠️ Top pitfall: Comparing the 90% and 80% in-range figures directly instead of computing the variance of X and Y: the two density tables feed the formula, and the single numbers 26.2 and 730.06 decide.

Self-check: Why does the old drug win when both drugs average 70 heartbeats per minute?

Connects to: Measures of Central Tendency

The Normal Distribution Preview

Must-know: Large samples will generally follow a normal distribution (small ones may not) — that is the law of large numbers; the normal distribution underpins estimations, hypothesis testing, and p-values.

⚠️ Top pitfall: Expecting every small data set to be bell-shaped; with only 26 students, one or two unusual values can make the sample lopsided, while 445 students tend toward the normal.

Self-check: Why is a 445-student class more likely to behave like a normal distribution than a 26-student class?

Connects to: Measures of Central Tendency, and Variance and Dispersion

Course Tools, Textbooks, and Evaluation

Must-know: What is taught is what is examined: exam questions are quoted by chapter, page, and problem number; the exam asks for your inference on software output, not hand computation.

⚠️ Top pitfall: Assuming a programming prerequisite or heavy calculus: the course covers R from zero, and only a bit of linear algebra is needed — no full calculus solving.

Self-check: Which textbook supplies about 60 to 70% of the course content, and what does it cover?

Exam Guidance Summary

Must-know: What is taught is what is examined: the delivered material is the boundary of the exam.

⚠️ Top pitfall: Planning for the makeup exam instead of the regular one: there is no difference between the two, and makeup time piles up work.

Self-check: How are exam questions quoted from the textbook?

Connects to: Course Tools, Textbooks, and Evaluation

Key Industry Applications

Must-know: Every concept of this session has a named industry home: honest visualization in health communication, variance in manufacturing and drug trials, and direct p-values from industry software.

⚠️ Top pitfall: Treating visualization as decoration: the statistician's job is to check whether the picture the numbers paint is the true picture before reports reach the public.

Self-check: Why does the lower-variance drug win the pharmaceutical trial?

Connects to: Statistical Literacy: How Statistics Can Mislead, and Variance and Dispersion

Was this lecture useful?

Loading comments…
🤖

BitsNotes AI Assistant

Subject Notes Assistant

Configure AI Chat

Choose how to access the chatbot
Have your own API key?

Switch to "Bring Your Own Key" tab above for unlimited access with any OpenAI-compatible provider.

🔑 Enter API key above to fetch live models from provider, or enter model name manually.
OpenAI-Compatible API Support

Choose any provider preset (Gemini, DeepSeek, Kimi, GLM, MiniMax, Qwen, OpenAI, Groq, Ollama, etc.) or enter a custom endpoint URL.

Security & Privacy First

Your API key is sent directly from your browser to your specified provider. BitsNotes servers never store or see your key.