Skip to main content

Inferential Statistics in Practice: From Probability to ANOVA


🔍 Project Overview 

This project demonstrates the application of inferential statistics to solve multiple real-world problems across sports analytics, manufacturing quality control, marketing operations and healthcare.

The objective was to move beyond descriptive statistics and apply probability theory, hypothesis testing, and ANOVA techniques to draw meaningful conclusions and support data-driven decision-making.

Download Complete Report from Git

Open on Git


🎯 Key Objectives

  • Apply probability concepts to real datasets

  • Use normal distribution and Z-tests for quality analysis

  • Perform hypothesis testing (Z-test, T-test)

  • Analyze multi-factor effects using One-Way & Two-Way ANOVA

  • Translate statistical results into business insights and recommendations


🧠 Problem 1: Sports Injury Probability Analysis

Business Question

Can player position help explain the likelihood of foot injuries in a football team?

Approach

  • Used conditional probability and joint probability

  • Analyzed injury distribution across playing positions

Key Insight

  • Overall injury probability: 61%

  • Strikers had the highest injury likelihood among injured players

  • Player position plays a significant role in injury risk

Impact

Helps coaching and medical staff focus preventive care strategies on high-risk positions.


🏭 Problem 2: Manufacturing Quality Control (Normal Distribution)

Business Question

What proportion of cement gunny bags fail strength requirements?

Approach

  • Assumed normal distribution

  • Used Z-score-based probability estimation

  • Visualized probability regions for decision clarity

Key Insights

  • ~11% of bags fall below minimum strength threshold

  • Over 82% meet acceptable strength criteria

  • Identified risk zones contributing to material loss

Impact

Supports supply chain quality checks and reduces wastage risk.


🧪 Problem 3: Stone Hardness Testing (Hypothesis Testing)

Business Question

Are unpolished stones suitable for high-quality printing?

Statistical Techniques Used

  • Z-test (large sample, known population mean)

  • Independent two-sample T-test

  • Outlier treatment and distribution analysis

Key Findings

  • Mean hardness of unpolished stones is significantly below required threshold

  • Polished stones show higher and more consistent hardness

Recommendation

Zingaro is justified in rejecting unpolished stones for printing applications.


🦷 Problem 4: Dental Implant Hardness Analysis (ANOVA)

Business Question

How do dentist, method, and alloy influence implant hardness?

Techniques Used

  • One-Way ANOVA

  • Two-Way ANOVA with interaction effects

  • Shapiro-Wilk Test (normality)

  • Levene Test (variance equality)

  • Tukey post-hoc analysis

Key Insights

  • Dentist alone does not significantly impact hardness

  • Implant method significantly affects hardness

  • Strong interaction exists between dentist and method

  • Optimal methods vary by alloy type

Business Impact

  • Standardizes implant procedures

  • Improves treatment outcomes

  • Reduces variability in medical results


🛠 Skills Demonstrated

Statistical & Analytical Skills

  • Probability theory

  • Hypothesis testing

  • Z-test, T-test

  • One-Way & Two-Way ANOVA

  • Post-hoc analysis

Tools & Techniques

  • Python

  • Pandas, NumPy

  • SciPy, StatsModels

  • Data visualization

  • Statistical interpretation


📈 Overall Impact

This project showcases the ability to:

  • Choose the right statistical test for each problem

  • Validate assumptions before modeling

  • Interpret statistical output in business terms

  • Support decisions with data-backed evidence


🏁 Conclusion

Inferential statistics is a critical foundation for data science and analytics.
This project demonstrates how statistical methods can directly support sports strategy, manufacturing quality, marketing optimization, and healthcare decision-making.













Comments

Popular posts from this blog

Using NLP for Text Analytics with HTML Links, Stop Words, and Sentiment Analysis in Python

  In the world of data science, text analytics plays a crucial role in deriving insights from large volumes of unstructured text data. Whether you're analyzing customer feedback, social media posts, or web articles, natural language processing (NLP) can help you extract meaningful information. One interesting challenge in text analysis involves handling HTML content, extracting meaningful text, and performing sentiment analysis based on predefined positive and negative word lists. In this blog post, we will dive into how to use Python and NLP techniques to analyze text data from HTML links, filter out stop words, and calculate various metrics such as positive/negative ratings, article length, and average sentence length. Prerequisites To follow along with the examples in this article, you need to have the following Python packages installed: requests (to fetch HTML content) beautifulsoup4 (for parsing HTML) nltk (for natural language processing tasks) re (for regular exp...

NumPy and Pandas for Data Science: A Comprehensive Guide

In the world of Data Science , working with large datasets, performing data manipulation, and analyzing numerical information is a fundamental task. To make these tasks easier and more efficient, Python has two powerful libraries: NumPy and Pandas . These libraries are widely used for data manipulation, analysis, and visualization and are crucial tools for any data scientist. Let’s take a deep dive into both NumPy and Pandas , exploring their functionality and how they empower data scientists to work smarter and faster. 1. What is NumPy? NumPy (Numerical Python) is an open-source library used for numerical computing in Python. It provides support for working with large, multi-dimensional arrays and matrices, and offers a wide range of mathematical functions to operate on these arrays. Key Features of NumPy: Efficient Array Operations: NumPy arrays, or ndarrays , are far more efficient in terms of memory and computational speed compared to Python’s native lists. Vectorizati...

Understanding Neural Network Models for Regression: ANN, RNN, and CNN

In the world of machine learning, neural networks play a crucial role in solving complex problems. They have shown remarkable performance in various domains, from image classification to natural language processing. However, one of the fundamental tasks that neural networks can perform is regression —predicting continuous values based on input features. In this blog post, we'll explore three types of neural network models— Artificial Neural Networks (ANN) , Recurrent Neural Networks (RNN) , and Convolutional Neural Networks (CNN) —and discuss how they can be used for regression tasks. Additionally, we'll walk through code examples and explain how to train these models for regression problems. What is Regression? Regression is a type of supervised learning where the model is trained to predict continuous values. Common examples of regression tasks include predicting house prices, stock market trends, or temperature forecasting. The primary goal is to find the best-fit line (...