Exam preparation
Possible questions across all units, grouped by topic and by marks (2, 5, and 10, in the SRM pattern). These are practice questions to guide revision.
Exam preparation: question bank
2 Marks
- Define statistics as a branch of mathematics, and name its five core activities.
- Distinguish a statistic from a mere numerical fact, giving one example of each.
- State what is meant by "aggregate of facts" as a characteristic of statistics.
- Define data, and state two forms that data can take other than numbers.
- Distinguish descriptive statistics from inferential statistics in one sentence each.
5 Marks
-
Classify each of the following as a genuine statistic or a mere number, and justify your answer in each case: (i) the temperature in Chennai at 9 a.m. today was 31 degrees Celsius; (ii) the average monthly rainfall in Chennai is 120 millimetres; (iii) Priya's salary is 60,000 rupees; (iv) the mean salary of employees in a firm is 55,000 rupees.
-
Explain the difference between descriptive and inferential statistics with a worked example. Take a survey of 20 students whose average mark is 65 out of a cohort of 500, and show clearly which statement about this data is descriptive and which is inferential, and why the inferential statement carries uncertainty.
-
A manager states that "the happiness of employees" is one of the company's key statistics. Explain, using the characteristic that a statistic must be numerically expressed and the characteristic that it must be enumerated or estimated, why this statement is problematic, and rewrite it so that it qualifies as a genuine statistic.
10 Marks
-
A smart farming company has deployed IoT sensors across 15 farms to monitor wheat production and records the crop yield in tonnes per hectare for each farm after harvest. The agricultural officer then reports the average yield across the 15 farms, and on the strength of that figure claims that farms of this type in the surrounding district can be expected to yield about the same amount next season. a) Identify the statistical activity being carried out when the officer records and averages the 15 yields, and name which of the five core activities of statistics are involved. b) Discuss which characteristics of statistics are at play in this scenario, referring specifically to aggregate of facts, affected by several causes, and comparability. c) Reason about which part of the officer's work is descriptive and which part is inferential, and explain why the claim about next season's district yield carries uncertainty.
-
An e-commerce company tracks the delivery time in hours of 18 customer orders from its Chennai warehouse. The logistics manager summarises these 18 times into a single reported "average delivery time", and on that basis argues to senior management that customers across the whole city receive their orders within roughly the same window. a) Identify the statistical activity involved in summarising the 18 delivery times, and state whether the reported average is a statistic or a mere number, justifying your answer. b) Explain the characteristics of statistics that apply here, including why delivery time is affected by several causes and why the data must be collected systematically and for a planned purpose to be trustworthy. c) Reason about whether the manager's argument to senior management is a descriptive or an inferential use of statistics, and explain what would have to be true of the 18 orders for the generalisation to the whole city to be justified.
2 Marks
- Define quantitative data and qualitative data, and give one example of each.
- Distinguish between a discrete numeric variable and a continuous numeric variable, with one example of each.
- Distinguish between nominal data and ordinal data. State the single question that separates them.
- Define primary data and secondary data, and state one risk of relying on secondary data.
- Define binary (dichotomous) data and state how it differs from nominal data.
5 Marks
- Classify each of the following variables by type, giving the full path on the variable-type tree (numeric or categorical, then the sub-type) and a one-line justification for each: number of defective items in a box; amount of rainfall in millimetres; blood group; education level (primary, secondary, high school, college, university); presence or absence of a disease.
- Explain the difference between structured and unstructured data. For each, give a concrete example and state why unstructured data is generally harder to analyse. Conclude with why most real-world data is unstructured.
- Explain the four classifications of data by structure over time: time series, cross-sectional, pooled, and panel. For each, give one example, and clearly state the two axes (one time point or many; same units or not) that distinguish pooled data from panel data.
10 Marks
-
A city transport authority stores one record per bus trip in the following table. The columns are: Trip ID (a unique number assigned to each trip), Route Number (an integer such as 12 or 47 identifying the route), Departure Time (the clock time of departure), Passengers Boarded (a whole-number count), Distance Travelled (kilometres, recorded to one decimal place), Payment Mode (coded 1 for cash, 2 for card, 3 for mobile wallet), and Service Rating (coded 1 for poor, 2 for average, 3 for good). Every trip in the table is recorded on the same single day. a) Classify every column by variable type, giving the full path on the variable-type tree and a justification for each. b) Identify the classification of this dataset by structure over time, and justify your answer using the two axes of time points and units. c) The Route Number, Payment Mode, and Service Rating columns all look numeric. For each, state whether arithmetic such as computing a mean is valid, and explain why or why not. d) State the general rule that decides whether a numeric-looking column is quantitative or categorical, and apply it to Payment Mode.
-
A public-health team surveys the same 200 households in a district once every year for five years. Each yearly record contains: Household ID (a fixed identifier reused across all five years), Survey Year (2021 to 2025), District Zone (coded North, South, East, West), Monthly Income (rupees, measured to the nearest rupee), Number of Children (a whole-number count), Access to Clean Water (coded 1 for yes, 0 for no), and Nutrition Status (coded 1 for undernourished, 2 for adequate, 3 for well-nourished). a) Classify every column by variable type, giving the full path on the variable-type tree and a justification for each. b) Identify the classification of this dataset by structure over time. Justify your choice, and explain precisely why it is panel data and not merely pooled data. c) Access to Clean Water is stored as 1 and 0. Explain which computations on this column are valid and which are meaningless, and justify your answer with reference to what the codes represent. d) The team proposes reporting the "average Nutrition Status" across all households. Explain whether this computation is statistically valid, and recommend a more appropriate summary for this column.
2 Marks
- Define a population and a sample, and state the relationship between them.
- What is a sampling frame? Give one example and explain how it differs from the population.
- Distinguish between sampling with replacement and sampling without replacement.
- Distinguish between probability sampling and non-probability sampling in terms of the chance of inclusion.
- Define sampling error, and name the two kinds of error it comprises.
5 Marks
- A quality engineer must estimate the average life of electric bulbs coming off a production line, where the test destroys each bulb. Recommend a suitable approach to data collection, state whether a full census is possible, and justify why sampling is required here rather than merely convenient.
- Explain stratified sampling and cluster sampling using the same population of 100 students who are to be sampled down to 20. For each method, describe how the 20 students are selected, and state clearly how the two methods differ in the way they treat groups and members.
- Distinguish random sampling error from non-random sampling error. For each, state its cause, whether it over- or under-estimates the true value, and whether increasing the sample size reduces it. Support your answer with the sources of non-random error.
10 Marks
-
A state health department wants to estimate the proportion of adults in a large city who have received a particular vaccine. The city has a known register of residential addresses grouped by administrative ward. Design a sampling plan for this study using each of the five probability methods, namely simple random, systematic, stratified, cluster, and multistage. For each method, identify the sampling frame, state the sample it would produce, and discuss the bias it may introduce and the situation in which it is the most appropriate choice for this study. Conclude by stating which method you would recommend and why.
-
A researcher studying the working conditions of gig-economy delivery riders in a city has no complete list of such riders, since they work for many platforms and turn over quickly. Consider the same target population and design a sampling plan using each of the four non-probability methods, namely convenience, purposive, snowball, and quota. For each method, state exactly how the sample would be recruited and the sample it produces, discuss the bias it introduces, and explain when that method would be appropriate. Finally, explain why non-probability sampling is used here rather than probability sampling, and state what limitation this places on any generalisation from the results.
2 Marks
- Define the median and state the rule for computing it when the number of observations is even.
- When is the mode preferred over the mean as a measure of central tendency? Give one situation.
- What is the trimmed mean, and why is the data sorted before it is computed?
- State the formula for the interquartile range and explain what quantity it measures.
- Write the two limits used by the 1.5 times IQR rule and state which values are flagged as outliers.
5 Marks
-
Given the dataset , find the mean, the median, and the mode. Show the sorting step for the median.
-
Given the dataset , find the mean and the mode, and identify whether the data is unimodal, bimodal, or multimodal.
-
Given the sorted dataset , compute , , the IQR, and the range using the positional formula.
10 Marks
-
A hospital records the waiting time (in minutes) of 15 patients at its outpatient department during one morning shift. The values are:
The administrator suspects that one unusually long wait has distorted the reported average and wants a clear summary of the morning's performance.
a) Calculate the mean waiting time.
b) Determine the median and the mode of the waiting times.
c) Calculate the first quartile , the third quartile , the interquartile range (IQR), and the range.
d) Using the 1.5 times IQR rule, identify any outlier in the data. Then recommend which measure of central tendency best represents the typical waiting time, and justify your choice with reference to the outlier.
-
An online retailer records the daily number of returned orders processed at its Chennai warehouse over 16 working days. The values are:
The operations manager wants to know whether the count on the last day is exceptional before setting a staffing target.
a) Calculate the mean number of returned orders.
b) Determine the median and the mode of the daily counts.
c) Calculate the first quartile , the third quartile , the interquartile range (IQR), and the range.
d) Using the 1.5 times IQR rule, identify any outlier in the data. Then recommend which measure of central tendency best represents a typical day, and justify your choice with reference to the outlier.
2 Marks
- Define variance and state the two formulas that distinguish the population case from the sample case.
- Why do we square the deviations from the mean when computing the variance instead of simply summing them?
- Explain why the coefficient of variation is unit-free, and state one situation in which this property makes it preferable to the standard deviation.
- State the difference between population variance and sample variance, including the divisor used in each and the reason for the difference.
- Define the range and give one reason why it is considered an unreliable measure of dispersion.
5 Marks
-
A machine produces rods whose lengths in centimetres are [8, 10, 12, 14, 16]. Treating the five rods as the entire population, compute the range, the population variance, and the population standard deviation. Show the sum of squared deviations.
-
The daily number of support tickets over a working week is [4, 6, 8, 10, 12]. Treating this as a sample, compute the sample variance and the sample standard deviation. Show the sum of squared deviations and state clearly which divisor you used and why.
-
Two batches of resistors are measured. Batch X has mean ohms and standard deviation ohms; Batch Y has mean ohms and standard deviation ohms. Compute the coefficient of variation for each batch and state which batch is more consistent in relative terms, explaining why the raw standard deviations alone would give a misleading answer.
10 Marks
-
Smart Agriculture, sensor reliability. A precision-farming company logs the soil-moisture readings (in percent) reported by two IoT sensors installed in the same field over five days. Sensor P reports [30, 32, 34, 33, 31] and Sensor Q reports [20, 26, 34, 40, 30]. The agronomist wants to know which sensor gives more stable readings before trusting one for irrigation decisions.
a) Calculate the range of each sensor's readings.
b) Treating each set of five readings as a sample, compute the sample variance and the sample standard deviation for each sensor. Show the sum of squared deviations from the mean in each case.
c) Calculate the coefficient of variation for each sensor.
d) Compare the consistency of the two sensors using the coefficient of variation and comment on which sensor is more consistent and should be trusted for irrigation decisions.
-
E-Commerce, delivery-time consistency. A logistics manager records the delivery times (in hours) for two courier partners handling the same route. Partner A delivers in [18, 20, 22, 19, 21] and Partner B delivers in [10, 30, 20, 25, 15]. The manager must recommend one partner based on how predictable the delivery times are.
a) Calculate the range of delivery times for each partner.
b) Treating each set of times as a sample, compute the sample variance and the sample standard deviation for each partner. Show the sum of squared deviations from the mean in each case.
c) Calculate the coefficient of variation for each partner.
d) Compare the consistency of the two partners using the coefficient of variation and comment on which partner offers more consistent delivery performance and should be recommended.
2 Marks
- Define skewness and state what a value of zero indicates about the distribution.
- Distinguish between positive skew and negative skew, naming the direction of the long tail in each case.
- What is a leptokurtic distribution? Write the formula for excess kurtosis in terms of the central moments and , state the sign it takes for a leptokurtic distribution, and describe its peak and tails.
- State the order of the mean, median, and mode for a right (positively) skewed distribution.
- Write Bowley's coefficient of skewness in terms of the quartiles , , and , and state the range of values it can take.
5 Marks
-
A dataset of monthly rainfall has first quartile , median , and third quartile . Compute Bowley's coefficient of skewness, and interpret the sign and magnitude of your result.
-
For a set of examination marks, the mean is , the median is , and the standard deviation is . Using the median form of Karl Pearson's coefficient, compute the coefficient of skewness and interpret what its sign implies about the position of the tail.
-
Explain the difference between skewness and kurtosis as descriptors of distribution shape. Define the three kurtosis regimes (platykurtic, mesokurtic, leptokurtic) with the sign of excess kurtosis for each, and explain why reported kurtosis normally places the normal distribution at zero.
10 Marks
-
A logistics firm records the delivery times (in hours) of orders from its Chennai warehouse. The summary statistics are: mean , median , mode , standard deviation , with quartiles , , and . One heavily delayed shipment took hours. a) Compute Karl Pearson's coefficient of skewness using the mode form, and Bowley's coefficient of skewness from the quartiles. b) State and interpret the direction and magnitude of skew indicated by each coefficient, and explain why the two coefficients differ in the presence of the delayed shipment. c) Describe the expected shape of the histogram of delivery times, and state whether you would expect the excess kurtosis to be positive or negative given the single large outlier. d) Explain what this shape implies for choosing the mean or the median to report the typical delivery time.
-
A university reports mid-semester marks for a Data Science course. The summary statistics are: mean , median , mode , standard deviation , with quartiles , , and . A few students scored well below the class, the lowest being . a) Compute Karl Pearson's coefficient of skewness using the median form, and Bowley's coefficient of skewness from the quartiles. b) State and interpret the direction and magnitude of skew, explaining which side carries the long tail. c) Describe the expected shape of the histogram of marks, and state whether the tail of low scorers would tend to make the distribution platykurtic or leptokurtic, giving your reasoning. d) Explain what this shape implies for choosing the mean or the median to summarise typical student performance, and justify your choice.
2 Marks
- Define covariance and state what its sign indicates about the relationship between two variables.
- Define the Pearson correlation coefficient and state what it measures beyond covariance.
- State the range of the correlation coefficient , and interpret the meaning of the values , , and .
- State two differences between covariance and correlation.
- Does a high correlation between two variables imply that one causes the other? Justify your answer in one sentence.
5 Marks
-
Explain the relationship between covariance and correlation. In your answer, write both formulas, state why correlation is called the normalised form of covariance, and explain why the two measures always share the same sign.
-
For the paired series compute the sample covariance using the divisor . Show the means, the deviation products, and the final value.
-
For the paired series compute the Pearson correlation coefficient , showing the covariance, both sample standard deviations, and the final value. Interpret the strength and direction of the result.
10 Marks
-
Digital Marketing: Advertising Spend and Sales. A retail company records its monthly advertising spend (in thousands of rupees) and the resulting sales (in lakhs of rupees) over six months:
Month Advertising spend Sales 1 2 4 2 4 6 3 6 9 4 8 10 5 10 11 6 12 14 The marketing manager wants to know whether higher advertising spend is associated with higher sales, and how strong that association is.
a) Compute the sample covariance using the divisor . Show the means and the deviation products. b) Compute the Pearson correlation coefficient , showing the two sample standard deviations and the arithmetic that combines them with the covariance. c) Interpret the strength and direction of the relationship in words. d) The manager concludes that increasing the advertising budget will directly cause sales to rise. Discuss whether this causal claim is justified from the correlation alone, and name one other factor that could explain the association.
-
IoT and Energy: Outdoor Temperature and Electricity Demand. A utility deploys sensors that log the average daily outdoor temperature (in degrees Celsius) and the household electricity demand (in kilowatt hours) on five summer days:
Day Temperature Demand 1 20 30 2 24 35 3 28 45 4 32 52 5 36 58 The analytics team wants to summarise how strongly demand tracks temperature.
a) Compute the sample covariance using the divisor . Show the means and the deviation products. b) Compute the Pearson correlation coefficient , showing both sample standard deviations and the final arithmetic. c) Interpret the strength and direction of the relationship, and state what the covariance alone could and could not have told the team. d) A colleague argues that the near-perfect correlation proves that higher temperature causes higher electricity demand. Discuss whether the causal claim is justified, and comment on whether a low correlation would have proved the two variables are independent.
2 Marks
-
State the fundamental (multiplication) counting principle for independent tasks and write the general formula for the total number of ways.
-
Define the factorial of a non-negative integer, and state the value of together with the reason the convention is adopted.
-
Define a permutation and a combination, and state in one sentence the rule that decides which to use for a given problem.
-
Write the formula for the number of distinct arrangements of objects when are alike of one kind, alike of another, and so on.
-
State the identity that relates and , and explain in words why the factor appears.
5 Marks
-
A four-character password is formed from the lowercase letters. Compute the number of possible passwords (a) when letters may be repeated and (b) when no letter may be repeated. State which counting rule each case uses.
-
Compute the number of distinct arrangements of all the letters of the word BALLOON (: two L, two O, and the rest distinct). Show the factorial expression and the arithmetic, and name the formula used.
-
Explain the difference between a permutation and a combination. Use the selection of objects from the distinct objects to illustrate: compute and , and verify numerically that .
10 Marks
-
A committee is to be formed from a pool of men and women, and a total of members must be chosen.
a) Compute the total number of ways to choose any members from the people, and name the counting rule used.
b) Compute the number of committees that contain exactly men and women, showing how the multiplication principle combines the two selections.
c) Compute the number of committees that contain at least women.
d) Assuming every committee is equally likely, compute the probability that a randomly formed committee has exactly men and women, and interpret the result.
-
A cricket-club locker room has distinct textbooks to be arranged in a single row on a shelf: are coaching manuals and are rulebooks.
a) Compute the total number of arrangements of all books with no restriction.
b) Compute the number of arrangements in which the rulebooks stay together, using the block method, and show the within-block ordering step.
c) Using the results of parts a) and b), compute the number of arrangements in which the rulebooks are NOT together.
d) A four-digit locker code is set using digits to . Compute the number of possible codes when digits may repeat and when they may not, and recommend which policy the club should adopt if it wants the larger number of possible codes, justifying the choice with the two counts.
2 Marks
- Define a random experiment and state its two defining features.
- Distinguish between an elementary event and a compound event, giving one example of each for a single roll of a die.
- State the three Kolmogorov axioms of probability.
- Define two mutually exclusive events and write the form the addition rule takes for them.
- Define two independent events and state the condition in words.
5 Marks
-
A standard six-sided die is rolled once. Let be the event "an even number" and be the event "a number greater than 4". a) Write out the sample space and the sets and . b) Compute , , and . c) Use the addition rule to compute . d) State whether and are mutually exclusive, and justify your answer.
-
Two cards are drawn from a standard 52-card deck without replacement. a) Compute the probability that both cards are kings. b) Compute the probability that the first card is a heart and the second card is a spade. c) Explain why the second draw is conditional on the first, and how the answers would change if the draws were made with replacement.
-
Explain, with reference to the three interpretations of probability (classical, empirical, subjective), how you would assign a probability in each of the following situations, and why the other two interpretations are less suitable: a) The probability of drawing the ace of spades from a well-shuffled deck. b) The probability that a particular manufacturing machine produces a defective part, given a log of 10,000 past parts of which 150 were defective. c) The probability that a newly founded company survives its first five years.
10 Marks
-
A quality-control bin contains 12 components: 7 are working and 5 are defective. An inspector draws two components at random without replacement and tests them. a) Compute the probability that both drawn components are defective. b) Compute the probability that both drawn components are working. c) Using the complement rule, compute the probability that at least one of the two drawn components is defective. d) A second inspector claims that "working on the first draw" and "defective on the second draw" are independent events. Determine whether this claim is correct by comparing with , and state a recommendation on whether the without-replacement scheme can be treated as independent.
-
A fair coin is tossed and, independently, a fair six-sided die is rolled. Let be the event "the coin shows heads" and let be the event "the die shows a 6". a) State the sample space of the combined experiment and its total number of equally likely outcomes. b) Compute , , and using the multiplication rule for independent events. c) Compute using the addition rule, and confirm the result by counting favourable outcomes directly. d) A student argues that because and can never be described together they must be mutually exclusive. Explain the error, decide whether and are mutually exclusive or independent, and recommend the mantra a student should use to avoid confusing the two ideas.
2 Marks
- Define marginal, joint, and conditional probability, and state how each is read off a two-way table.
- Write the formula for the conditional probability and state the condition under which it is defined.
- State the two equivalent conditions for events and to be independent.
- Define the terms prior, likelihood, and posterior as they appear in Bayes theorem.
- Distinguish sensitivity and specificity of a diagnostic test, writing each as a conditional probability.
5 Marks
-
A survey of people records exercise habit against high blood pressure:
High BP Normal BP Total Exercises 60 440 500 Does not 190 310 500 Total 250 750 1000 a) Compute the marginal probability and the joint probability . b) Compute and , and comment on why they differ. c) Determine numerically whether exercise and high blood pressure are independent.
-
In a factory, machine A produces of items and machine B produces the remaining . Machine A has a defect rate of and machine B a defect rate of . Using the law of total probability, compute the overall probability that a randomly selected item is defective. Then, given that an item is defective, use Bayes theorem to find the probability it was produced by machine B. Interpret the result.
-
A screening test for a condition has sensitivity and specificity . The condition affects of the tested population. A patient tests positive. Compute using the law of total probability and then the posterior probability using Bayes theorem. State in one sentence why the posterior is far below the sensitivity.
10 Marks
-
A hospital screens for a rare disease that affects of the population it serves. The available test has sensitivity and specificity . Consider a cohort of people drawn from this population.
a) Using natural frequencies, fill in the counts of true positives, false negatives, false positives, and true negatives across the cohort. b) Compute the total number of positive test results, and from it the posterior probability that a person who tests positive actually has the disease. c) Recompute the posterior directly from Bayes theorem using the probabilities, and confirm it matches the frequency answer. d) The prior prevalence rises to in a high-risk subgroup. Recompute for this subgroup and explain, with reference to the base-rate effect, why the same test is now far more informative. Recommend a testing policy for the general population.
-
An email service classifies messages as spam or legitimate. Overall, of incoming messages are spam. The word "offer" appears in of spam messages and in of legitimate messages.
a) Using the law of total probability, compute . b) Using Bayes theorem, compute the posterior probability that a message containing "offer" is spam. c) A second, independent word "winner" appears in of spam and of legitimate messages. Assuming the two words occur independently within each class, compute the posterior probability that a message containing both "offer" and "winner" is spam. d) Compare the single-word and two-word posteriors, explain how the likelihood ratio drives the update, and recommend whether the filter should flag on a single word or require corroborating evidence. Contrast this with the rare-disease case where the prior dominates.
2 Marks
- Define a discrete random variable and state the two conditions its probability mass function must satisfy.
- Write the probability mass function of the binomial distribution and state its mean and variance in terms of and .
- Write the probability mass function of the Poisson distribution and state its mean and variance in terms of .
- Distinguish a Bernoulli trial from a binomial experiment, and give the mean and variance of a single Bernoulli trial.
- State the condition under which the Poisson distribution approximates the binomial, and give the value of used in that approximation.
5 Marks
-
A discrete random variable takes the values with probabilities respectively. Verify that this is a valid probability mass function, then compute and using . Interpret the mean.
-
A fair coin is tossed 8 times. Let be the number of heads, so . a) Compute . b) Compute using the complement. c) State the mean and variance of .
-
A web server receives requests at an average rate of per second, modelled as Poisson. Let be the number of requests in a given second. a) Compute and . b) Compute . c) State the mean and variance of and explain what their equality tells you about the model.
10 Marks
-
A microchip fabrication line produces chips that are defective independently with probability . A quality inspector draws a sample of chips and lets be the number of defectives in the sample. a) State the distribution of and its four assumptions, and confirm they are met here. b) Compute the mean and variance of . c) Compute , the probability the sample is defect free, and , the probability of at least one defective. d) The line is judged acceptable if the expected number of defectives per 50-chip sample is at most 2. State whether the line is acceptable and justify the recommendation from your computed mean.
-
A hospital emergency department receives patients at an average rate of per hour, and arrivals are independent and occur at a constant rate. Let be the number of arrivals in a given hour. a) State the distribution of and why the Poisson is the appropriate model rather than the binomial. b) Compute the mean and variance of . c) Compute , the probability of exactly five arrivals, and , the probability of at most two arrivals. d) Staffing is set so that the department can comfortably handle up to 3 arrivals per hour. Using from part c) as an indicator, comment on how often the department is likely to be understaffed and give a recommendation.
2 Marks
- Define a continuous random variable and explain why for such a variable.
- State what a probability density function represents and give the value of the total area under any valid PDF.
- Write the notation in words and state what each of and controls about the shape of the curve.
- Define the standard normal distribution and state the mean and standard deviation that characterise it.
- State the empirical rule, giving the approximate percentage of values within 1, 2, and 3 standard deviations of the mean.
5 Marks
-
Explain, with a labelled sketch of the bell curve, how a z-table returns the area to the left of a z-score, and describe how you would obtain (a) a right-tail probability and (b) an interval probability from such left-tail values.
-
The lifetime of a certain battery is normally distributed with mean hours and standard deviation hours. a) Standardise the value hours to a z-score. b) Using , find and interpret the result in context.
-
The daily sales of a shop are normally distributed with mean units and standard deviation units. a) Find the z-scores for 175 units and 250 units. b) Using and , compute and state what proportion of days fall in this range.
10 Marks
-
A university reports that student marks in a statistics module are normally distributed with mean and standard deviation . Use the z-table values , , and . a) A student scores 55. Compute the z-score and find , the proportion of students scoring below this mark. b) Find , the proportion of students scoring above 75. c) Find , the proportion scoring between 55 and 75, and verify that your three regions from parts a) to c) sum to approximately 1. d) The top 10% of students receive a distinction. Find the minimum mark required, using , and interpret the result.
-
A bottling plant fills bottles with a volume that is normally distributed with mean ml and standard deviation ml. A bottle is rejected if it contains less than 494 ml or more than 508 ml. Use the z-table values , , and . a) Compute the z-scores for the two reject limits, 494 ml and 508 ml. b) Find the proportion of bottles rejected for being underfilled, , and the proportion rejected for being overfilled, . c) Find the proportion of bottles that are accepted, . d) The plant wants only the lowest 5% of volumes to fall below a warning threshold. Find that threshold volume using , and recommend whether the current lower reject limit of 494 ml is set appropriately.
Sampling Distribution and the Central Limit Theorem: Exam Questions
2 Marks
- Define a population parameter and a sample statistic, and give one example of each with its correct notation.
- State the formula for the standard error of the mean and explain in one sentence how it differs from the population standard deviation .
- State the central limit theorem, including the approximate distribution of for large .
- Explain the rule of thumb, and state what happens to the sampling distribution of when the population is itself normal.
- Write the formula for the z-score of a sample mean and identify what each symbol in the denominator represents.
5 Marks
-
A population has standard deviation . Compute the standard error of the sample mean for , , and . Comment on how the standard error changes as is quadrupled, and state the factor by which it changes.
-
Scores on an aptitude test have population mean and standard deviation . A sample of candidates is drawn.
- a) Compute the standard error of the sample mean.
- b) Compute , showing the z-score and the tail area. Use .
- c) Interpret the result in one sentence.
-
The daily number of support tickets at a helpdesk has mean and standard deviation , and the distribution is right-skewed. A sample of days is recorded.
- a) State why the sample mean can still be treated as approximately normal despite the skew.
- b) Compute the standard error.
- c) Compute . Use and .
10 Marks
-
A bottling plant fills bottles whose contents have population mean millilitres and standard deviation millilitres. The filling distribution is not assumed normal. Quality control draws a sample of bottles and records the sample mean.
- a) Compute the standard error of the sample mean, and state the approximate distribution of under the central limit theorem.
- b) Compute . Use .
- c) Compute . Use and .
- d) Management wants to halve the standard error you found in part a). State the sample size required and briefly justify it using the relationship. Give one practical recommendation about the trade-off between precision and sampling effort.
-
A logistics firm measures delivery times in hours. The population is heavily right-skewed with mean hours and standard deviation hours. An analyst samples deliveries.
- a) Explain whether the sample mean can be treated as approximately normal, referring to both the population shape and the sample size, and compute the standard error.
- b) Compute . Use .
- c) The analyst writes a short numpy simulation that draws 20000 samples of size 100 from an exponential population with mean 8 and records each sample mean. State the approximate value of the mean of those 20000 sample means and the approximate value of their standard deviation, and name the two central limit theorem facts each result confirms.
- d) The firm claims its average delivery time is under 8 hours. Given the sampling distribution you described, explain how a single observed sample mean could be used to assess this claim, and recommend whether a larger sample would sharpen the assessment.
2 Marks
-
Distinguish between an estimator and an estimate, giving one example of each for the population mean.
-
State the three desirable properties of a good point estimator and name them precisely.
-
Define the margin of error for a confidence interval and write its formula for the case when the population standard deviation is known.
-
State the correct interpretation of a "95% confidence interval" and give one common misinterpretation that must be avoided.
-
State the two conditions under which you would use the t distribution rather than the z distribution when building a confidence interval for a population mean.
5 Marks
-
Explain why the sample variance is defined with a divisor of rather than . In your answer, connect the choice to the property of unbiasedness and explain what would happen to the estimate of the population variance if were used instead. Then explain how an estimator can be biased yet still consistent.
-
A quality engineer measures the fill weight of a sample of bottles from a filling line. The sample mean is ml, and the process standard deviation is known from long experience to be ml. Construct a 95% confidence interval for the true mean fill weight. Show the standard error, the critical value, the margin of error, and the final interval, then interpret the interval in one sentence. (Use .)
-
A researcher records the drying times, in minutes, of a new paint for a sample of test panels. The sample mean is minutes and the sample standard deviation is minutes; is unknown. Construct a 95% confidence interval for the true mean drying time. State the degrees of freedom, show the standard error, the critical value, the margin of error, and the final interval, and interpret it. (Use .)
10 Marks
-
A hospital laboratory is studying the fasting blood glucose level of adult patients, measured in mg/dL. The measuring instrument is well characterised, so the population standard deviation is taken as known at mg/dL. A random sample of patients gives a sample mean of mg/dL.
a) State which distribution (z or t) is appropriate here and justify the choice from the information given.
b) Compute the standard error of the mean and construct a 95% confidence interval for the true mean glucose level. (Use .)
c) Recompute the interval at the 99% confidence level. (Use .) Compare the two interval widths and explain, in terms of the critical value, why one is wider.
d) The laboratory wants a 95% interval with a margin of error no larger than mg/dL. Using , determine the minimum sample size required, and state a one-sentence recommendation to the laboratory.
-
A logistics company is auditing the delivery time, in hours, for parcels sent to a particular region. Historical spread is not reliable, so the population standard deviation is unknown. A random sample of deliveries gives a sample mean of hours and a sample standard deviation of hours. The delivery times are believed to be approximately normally distributed.
a) State which distribution (z or t) is appropriate here and justify the choice, including the number of degrees of freedom.
b) Compute the standard error of the mean and construct a 95% confidence interval for the true mean delivery time. (Use .)
c) The company advertises an average delivery time of "no more than 28 hours." Using your interval from part b), state whether the data are consistent with that advertised claim, and justify your answer from the interval.
d) Explain how the interval would change if the sample size were increased to while the sample mean and sample standard deviation stayed the same, and make a recommendation about whether collecting more data is worthwhile here. Address both the critical value and the standard error in your explanation.
2 Marks
-
Define a statistical hypothesis and explain why the null hypothesis must always be stated with an equality.
-
Distinguish the null hypothesis from the alternative hypothesis , giving the sign each may carry.
-
State the difference between a one-tailed and a two-tailed test, and explain which hypothesis fixes the type.
-
Define the p-value precisely, and state clearly what it is not.
-
Explain why the correct conclusion of a non-significant test is "fail to reject " rather than "accept ".
5 Marks
-
A researcher wants to know whether a coaching class changes the average score on a standard test whose long-run mean is 60. Write the hypotheses and , state whether the test is one-tailed or two-tailed, and explain how the wording of the research question determined your choice. Then describe, in the correct order, the six steps you would follow to carry out the test.
-
A production line is designed to fill packets to a mean weight of g with a known population standard deviation g. A random sample of packets has mean g. Test at whether the mean fill weight differs from 250 g. Formulate the hypotheses, compute the z-test statistic, obtain the two-tailed p-value, and state the decision with the critical values . (Use .)
-
A battery maker claims a mean life of at least 40 hours. A quality analyst suspects the true mean is lower and draws batteries with mean hours; the population standard deviation is known to be hours. Test at using a left-tailed z-test. Formulate the hypotheses, compute the test statistic, obtain the p-value, compare with the critical value , and state the decision. (Use .)
10 Marks
-
A bakery states that its loaves weigh 400 g on average. A consumer group, suspecting the loaves are underweight, weighs a random sample of loaves and finds a sample mean of g. The population standard deviation is known to be g, and the weights are approximately normal. Working at :
a) State the null and alternative hypotheses and explain why this is a left-tailed test.
b) Compute the standard error and the z-test statistic, showing the arithmetic.
c) Using the critical value and the p-value (take ), reach a decision under both the critical-value and the p-value approaches, and confirm that they agree.
d) State the conclusion in context, and explain what would change if the group had instead only asked whether the mean weight differs from 400 g.
-
A call centre reports that the mean handling time per call is 300 seconds. After a new training programme, the manager wants to know whether the mean handling time has changed in either direction. A random sample of calls after training has a mean of seconds; the population standard deviation is known to be seconds. Working at :
a) State the hypotheses and identify the type of test, justifying your choice from the manager's question.
b) Compute the standard error and the z-test statistic, showing the arithmetic.
c) Obtain the two-tailed p-value (take ) and compare the statistic with the critical values , reaching a decision.
d) State the conclusion in context, write the scipy expression you would use to obtain the p-value from the statistic, and explain why the manager should not phrase the outcome as "accepting" any hypothesis.
2 Marks
- Define a Type I error and state its probability.
- Define a Type II error and state its probability.
- Define the power of a test and express it in terms of .
- State whether and are complements, and justify your answer in one sentence.
- State the rule of thumb for choosing between a z-test and a t-test for a single population mean.
5 Marks
-
A quality inspector tests "a batch of components is within specification" against "the batch is defective". For each of the following outcomes, identify whether it is a Type I error, a Type II error, or a correct decision, and name the probability associated with each error. a) The batch is within specification, but the inspector rejects it. b) The batch is defective, and the inspector rejects it. c) The batch is defective, but the inspector accepts it. d) State which single lever the inspector can change to reduce both error probabilities together, and why it works.
-
A one-sample z-test uses against with , , and . The true mean is . a) Compute the standard error of the mean. b) Find the critical value of above which is rejected (use ). c) Compute , the probability of a Type II error, using . d) State the power of the test.
-
For a fixed effect size, explain the effect on power of each of the following changes, and state the cost (if any) of each: a) decreasing ; b) increasing the sample size ; c) switching from a two-tailed to a one-tailed test in the correct direction.
10 Marks
-
A hospital claims that the mean recovery time after a standard procedure is 7 days. A researcher suspects the new protocol changes recovery time and records a sample of patients with sample mean days and sample standard deviation days. The population standard deviation is unknown and recovery times are approximately normal. Use . a) Using the decision guide (parameter type, number of samples, whether is known, sample size), select the correct test statistic and justify each choice. b) State and , and state whether the test is one-tailed or two-tailed given the researcher's suspicion. c) Compute the test statistic. The critical value is . State the decision. d) Suppose the researcher had instead claimed in advance that the new protocol only shortens recovery time. Explain how this changes the number of tails, the critical value, and the power of the test, and comment on whether choosing this direction after seeing would be legitimate.
-
A marketing team runs two versions of a landing page. Version A is shown to visitors and 60 sign up; version B is shown to visitors and 84 sign up. The team wants to know whether the sign-up proportions differ. Use and . a) Using the decision guide, identify the correct test statistic and justify why a z-test is appropriate for this parameter. b) Compute the two sample proportions and . c) The team originally planned to test only 40 visitors per version. Explain, in terms of the standard error and the overlap of the sampling distributions, why the larger sample of 400 per version gives the test more power to detect a real difference. d) The team lead argues that they should lower to 0.01 to be "more certain". Explain what this does to the Type I error rate, the Type II error rate, and the power of the test, and recommend whether this is a good idea given their goal of detecting a genuine difference.