Business Statistics II: Inference, Regression, and Decision Making
❦
Sampling Distributions and Estimation
- The sampling distribution of the mean has mean μ and standard error σ ÷ √n. The central limit theorem says that for large n the distribution of x̄ is approximately normal, whatever the population distribution.
- Confidence interval for a mean (σ known or large n): x̄ ± z × σ/√n, with z = 1.96 for 95 percent confidence and 2.58 for 99 percent. For small samples with unknown σ, use the t-distribution.
- Confidence interval for a proportion: p̂ ± z√[p̂(1 − p̂)/n].
- Sample size for estimating a mean: n = (zσ/E)², where E is the allowable error.
- Example: a sample of 100 customers has a mean spend of 2 000 and σ = 400. The 95 percent interval = 2 000 ± 1.96 × 40 = 2 000 ± 78.4 (1 921.6 to 2 078.4).
Hypothesis Testing
1. State the null hypothesis (H₀) and alternative (H₁).
2. Choose the significance level α (commonly 0.05).
3. Compute the test statistic: for a mean z = (x̄ − μ₀) ÷ (σ/√n).
4. Compare with the critical value (for a two-tailed test at 5 percent, ±1.96), or compute the p-value.
5. Decide and state the conclusion in business terms.
Type I error (rejecting a true H₀; probability α) and Type II error (failing to reject a false H₀). Tests include the z-test and t-test for means, tests for proportions, the chi-square test for independence and goodness of fit (χ² = Σ(O − E)² ÷ E), and, in some syllabi, ANOVA. A p-value is not the probability that the null hypothesis is true, and statistical significance is not the same as practical importance (Wasserstein & Lazar, 2016).
Correlation and Regression
- Scatter diagrams show relationships; the Pearson correlation coefficient r ranges from −1 to +1. Spearman's rank correlation is used for ranked data. Correlation does not prove causation.
- Simple linear regression: ŷ = a + bx, where b = [nΣxy − ΣxΣy] ÷ [nΣx² − (Σx)²] and a = ȳ − b x̄. The coefficient of determination R² is the proportion of variation in y explained by x.
- Worked example. Advertising spend x (in thousands): 1, 2, 3, 4, 5; sales y (in thousands): 3, 5, 6, 8, 10. Σx = 15, Σy = 32, Σxy = 113, Σx² = 55, Σy² = 234, n = 5. b = (5 × 113 − 15 × 32) ÷ (5 × 55 − 225) = 85 ÷ 50 = 1.7; a = 6.4 − 1.7 × 3 = 1.3. The regression line is ŷ = 1.3 + 1.7x. For x = 6, the predicted sales = 11.5. The correlation r = 85 ÷ √(50 × 146) ≈ 0.995, so R² ≈ 0.99.
- Multiple regression: ŷ = a + b₁x₁ + b₂x₂ + …; interpret each coefficient as the change in y for a unit change in that variable, holding the others constant. Check assumptions and avoid extrapolation beyond the data; beware of multicollinearity.
Forecasting
Methods include trend projection, moving averages, exponential smoothing, and regression-based forecasting. Evaluate accuracy by the mean absolute deviation (MAD) or the mean squared error (MSE).
Decision Making Under Uncertainty
- Decision tables (payoff matrices): criteria such as maximin (pessimistic), maximax (optimistic), minimax regret, and expected monetary value (EMV) = Σ(probability × payoff).
- Decision trees: choose the option with the best EMV, working backward.
- Value of information.
- Example: a project yields 200 000 with probability 0.6 and −50 000 with probability 0.4. EMV = 0.6 × 200 000 + 0.4 × (−50 000) = 120 000 − 20 000 = 100 000.
Statistical Software and Ethics
Use spreadsheets (Excel or Google Sheets) or statistical software (R, SPSS, or Python) to compute results. Report honestly: describe the data and methods, avoid selecting only favourable results, and respect privacy.
Common Mistakes
- Interpreting a confidence interval as the range containing 95 percent of the data.
- Confusing correlation with causation.
- Predicting outside the range of the data.
CHAPTER 18