Types of Statistics: Descriptive and Inferential Statistics

Jasmine Grover logo

Jasmine Grover

Education Journalist | Study Abroad Lead

Statistics is a branch of mathematics that involves the collection, analysis, interpretation, presentation, and organization of data. It is a powerful tool that enables us to make informed decisions based on data-driven insights. 

  • Statistics is widely used in various fields, including business, healthcare, engineering, social sciences, economics, and many others.
  • There are two types of statistics - descriptive and inferential statistics.
  • Descriptive statistics summarizes and describes data using measures such as mean, median, mode, range, variance, and standard deviation.
  • Inferential statistics involves making inferences and predictions about a population based on a sample. 
  • The main objective of statistics is to describe, summarize, and draw inferences from data. 
  • To achieve this, statisticians use various techniques such as sampling, data visualization, regression analysis, hypothesis testing, and statistical inference. 
  • Through these methods, statisticians are able to uncover patterns, trends, and relationships within datasets and provide insights that can help organizations make informed decisions.

Key Terms: Statistics, Data, Descriptive Statistics, Inferential Statistics, Mean, Mode, Median, Range, Standard Deviation, Variance.


What is Statistics?

[Click Here for Sample Questions]

Statistics is the process of collecting, analyzing, interpreting, and presenting data to gain insights into various phenomena. 

  • It is a powerful tool that helps us understand the world around us by providing a framework for organizing and making sense of complex information.
  • Statistics can be used to describe and summarize data, identify patterns and relationships, and make predictions based on past observations. 
  • By using statistical techniques, we can gain insights into a wide range of phenomena, from the behaviour of financial markets to the spread of diseases and the impact of social policies. 
  • Statistics allows us to identify trends, make predictions, and make informed decisions based on data-driven insights.

Also Read:


Types of Statistics

[Click Here for Sample Questions]

Statistics can be broadly classified into two main types: descriptive statistics and inferential statistics.

Descriptive statistics involves the collection, organization, analysis, and presentation of data in a way that describes its main features. 

  • This type of statistics focuses on summarizing and describing the characteristics of a dataset, including measures such as mean, median, mode, standard deviation, and variance. 
  • Descriptive statistics can be used to provide insights into trends, patterns, and relationships within a dataset.

Inferential statistics involves making inferences and drawing conclusions about a larger population based on a sample of data. 

  • This type of statistics uses probability theory and statistical inference to make predictions about a larger group based on a smaller subset of that group. 
  • Inferential statistics can be used to test hypotheses, make predictions, and estimate population parameters.

Some other types of statistics include

  • Exploratory data analysis: It involves using various visualization techniques to understand the characteristics of a dataset and identify patterns and relationships within it.
  • Regression analysis: This is a statistical method used to examine the relationship between two or more variables, and to model and predict their behavior.
  • Time series analysis: This involves analyzing data that is collected over time, and identifying trends and patterns that occur over different time intervals.
  • Multivariate analysis: This is a statistical technique used to analyze multiple variables at the same time, and to examine the relationships between them.

Descriptive statistics

Descriptive statistics is a type of statistics that involves summarizing and describing the main features of a dataset. This can include measures such as the mean, median, mode, standard deviation, and variance. Here are some examples of descriptive statistics:

  • Mean: The mean is the average value of a dataset. 

Example: For a dataset of test scores that includes scores of 70, 80, 90, and 100, the mean score would be (70+80+90+100)/4 = 85.

  • Median: The median is the middle value of a dataset when it is arranged in order. 

Example: For a dataset of salaries that includes values of $30,000, $40,000, $50,000, $60,000, and $70,000, the median salary would be $50,000.

  • Mode: The mode is the most common value in a dataset. 

Example: For a dataset of favourite colours that includes values of red, blue, green, green, yellow, and purple, the mode colour would be green.

  • Standard deviation: The standard deviation is a measure of the spread of a dataset. 

Example: For a dataset of ages that includes values of 20, 25, 30, 35, and 40, the standard deviation would be approximately 8.37.

  • Variance: The variance is another measure of the spread of a dataset. It is calculated by squaring the standard deviation. 

Example: If the standard deviation of a dataset is 4, the variance would be 16.

Inferential statistics

Inferential statistics is a type of statistics that involves making inferences and drawing conclusions about a larger population based on a sample of data. This type of statistics uses probability theory and statistical inference to make predictions about a larger group based on a smaller subset of that group. Here are some examples of inferential statistics:

Hypothesis testing: Hypothesis testing involves making a hypothesis about a population parameter (such as the mean or proportion) and then using a sample of data to test the hypothesis.

Example: A company may want to know if a new marketing campaign has led to an increase in sales. They could use inferential statistics to test the hypothesis that the mean sales have increased after the campaign, using a sample of sales data.

Confidence intervals: Confidence intervals are used to estimate the range of values that a population parameter is likely to fall within, based on a sample of data. 

Example: A researcher may want to estimate the average height of all adults in a city. They could use inferential statistics to calculate a confidence interval for the population mean height, based on a sample of heights.

Regression analysis: Regression analysis is a statistical technique that is used to examine the relationship between two or more variables, and to model and predict their behavior. 

Example: A company may want to know how changes in the price of their product affect sales. They could use inferential statistics to perform a regression analysis on sales data and price data, in order to estimate the relationship between the two variables.

Sampling: Sampling involves selecting a subset of a population in order to make inferences about the entire population. 

Example: A polling organization may want to estimate the percentage of voters who support a particular candidate in an upcoming election. They could use inferential statistics to select a representative sample of voters and then use the results of the sample to make inferences about the entire population of voters.


Types of Data in Statistics

[Click Here for Sample Questions]

In statistics, data is typically classified into four types: nominal, ordinal, interval, and ratio. These types of data have different properties and characteristics, and the choice of which type to use depends on the nature of the data and the research question being addressed.

Nominal data: Nominal data is categorical data that cannot be ordered or ranked. Examples of nominal data include gender, ethnicity, and job title. Nominal data is typically analyzed using frequency counts and percentages.

Ordinal data: Ordinal data is categorical data that can be ordered or ranked. Examples of ordinal data include survey responses such as "strongly disagree," "disagree," "neutral," "agree," and "strongly agree." Ordinal data is typically analyzed using median or mode.

Interval data: Interval data is numerical data where the distance between any two values is meaningful and equal, but there is no true zero point. Examples of interval data include temperature in Celsius or Fahrenheit, where zero does not represent the absence of temperature. Interval data is typically analyzed using mean or standard deviation.

Ratio data: Ratio data is numerical data where there is a true zero point, and the distance between any two values is meaningful and equal. Examples of ratio data include height, weight, and time. Ratio data is typically analyzed using mean, standard deviation, and coefficient of variation.


Types of variables in statistics

[Click Here for Sample Questions]

Variables are attributes or characteristics that can be measured and can vary from one observation to another. The four main types of variables in statistics are:

Nominal variables: Nominal variables are categorical variables that represent different categories or groups. Nominal variables cannot be ordered or ranked, and the categories have no inherent numerical value. Examples of nominal variables include gender, race, and religion.

Ordinal variables: Ordinal variables are categorical variables that have a natural order or ranking. The categories of ordinal variables can be ordered based on their relative position or importance. Examples of ordinal variables include education level, income bracket, and survey responses with the options of "strongly agree," "agree," "disagree," and "strongly disagree."

Interval variables: Interval variables are numeric variables that have equal intervals between values but do not have a true zero point. Interval variables are measured on a scale with meaningful units, but the zero point does not represent the absence of the variable being measured. Examples of interval variables include temperature in Celsius or Fahrenheit, and dates on the calendar.

Ratio variables: Ratio variables are numeric variables that have a true zero point and equal intervals between values. Ratio variables have meaningful units, and the zero point represents the complete absence of the variable being measured. Examples of ratio variables include weight, height, and time.

Also Read:


Measure of central tendency and dispersion

[Click Here for Sample Questions]

Measures of central tendency, such as the mean, median, and mode, describe the central or typical value of a set of data. 

  • Measures of dispersion, such as the range, variance, and standard deviation, describe the spread or variability of the data. 
  • They provide information about how widely the data is spread out.
  • Together, measures of central tendency and dispersion provide a more complete picture of a set of data. 

For example, two sets of data may have the same mean, but one set may have a larger range or variance, indicating that the data is more spread out than the other set.


Stages of statistics

[Click Here for Sample Questions]

There are generally four stages of statistics that are used to analyze data and make inferences about populations:

Data collection

The first stage of statistics involves collecting data. 

  • This can be done using various methods such as surveys, experiments, observations, or sampling. 
  • The data collected should be relevant to the research question and must be representative of the population being studied.

Data analysis

The second stage of statistics involves analyzing the data that has been collected. 

  • This includes organizing and summarizing the data using descriptive statistics such as measures of central tendency and variability. 
  • It also involves using inferential statistics to make inferences about the population based on the sample data.

Interpretation

The third stage of statistics involves interpreting the results obtained from the data analysis. 

  • This requires an understanding of the statistical methods used and their limitations, as well as an understanding of the research question being addressed. 
  • The interpretation of the results should be done in the context of the research question and the limitations of the study.

Communication

The final stage of statistics involves communicating the results of the data analysis and interpretation. 

  • This can be done through written reports, visual presentations, or oral presentations. 
  • The results should be communicated clearly and accurately to the intended audience and should include any limitations or assumptions made during the analysis.

Solved Example

Example 1: The following set of data representing the weights of 10 students in kilograms:

{50, 53, 55, 56, 57, 59, 60, 62, 65, 70}

Find the measures of central tendency and measure of dispersion

Solution: The mean weight is calculated by summing all the weights and dividing by the number of students:

(50 + 53 + 55 + 56 + 57 + 59 + 60 + 62 + 65 + 70) / 10 = 59.7 kg

Median: Since we have an even number of values, we take the average of the two middle values:

(57 + 59) / 2 = 58 kg

Mode: The mode weight is the most frequently occurring value, which in this case is 55 kg.

Range: The range is the difference between the highest and lowest values:

70 - 50 = 20 kg

Variance: It is calculated by taking the average of the squared differences between each value and the mean:

((50-59.7)2 + (53-59.7)2 + ... + (70-59.7)2) / 10 = 74.61

Standard deviation: √74.61 = 8.64 kg

Example 2: The following set of data representing the ages of 8 people

{22, 25, 27, 29, 31, 31, 32, 35}

Find the measures of central tendency and measure of dispersion

Solution: Mean: The mean age is calculated by summing all the ages and dividing by the number of people:

(22 + 25 + 27 + 29 + 31 + 31 + 32 + 35) / 8 = 29.5

Median: The median age is the middle value when the data is arranged in order from lowest to highest:

(29 + 31) / 2 = 30

Mode: The mode age is the most frequently occurring value, which in this case is 31.

Range: The range is the difference between the highest and lowest values:

35 - 22 = 13

Variance: The variance measures how spread out the data is from the mean. It is calculated by taking the average of the squared differences between each value and the mean:

((22-29.5)2 + (25-29.5)2 + ... + (35-29.5)2) / 8 = 21.75

Standard deviation: √21.75 = 4.66


Uses of statistics

[Click Here for Sample Questions]

Statistics are used in research to collect and analyze data, and to draw conclusions from the data.

  • In business, statistics are used for market research, forecasting, and analyzing financial data.
  • In healthcare, statistics are used to analyze patient data, assess treatment outcomes, and track disease trends.
  • In sports, statistics are used to evaluate player performance and to make strategic decisions.
  • Government agencies use statistics to analyze census data, track economic indicators, and monitor public health trends.
  • Statistics are also used in education to analyze student performance data and to evaluate teaching methods.

Importance of statistics

[Click Here for Sample Questions]

Statistics provides a way to collect and analyze data in a systematic and objective manner, which helps to reduce bias and increase accuracy in decision-making.

  • Statistics help us to understand patterns and trends in data, which can lead to new discoveries and innovations.
  • By using statistical methods, we can make predictions and forecasts based on historical data, which can be useful for planning and resource allocation.

Also Read:


Things to Remember

  1. Statistics is a branch of mathematics that involves collecting, analyzing, and interpreting data.
  2. Descriptive statistics and inferential statistics are the two main types of statistics.
  3. Applied statistics, mathematical statistics, Bayesian statistics, and nonparametric statistics are other types of statistics.
  4. Mean is the average value of the data set.
  5. Median is the middle value when the data set is ordered from lowest to highest.
  6. Mode is the most frequently occurring value in the data set.
  7. Standard deviation is the square root of the variance and is a commonly used measure of dispersion.

Sample Questions

Ques. What is the difference between descriptive and inferential statistics? (3 marks)

Ans. Descriptive statistics involve analyzing and summarizing data, such as calculating measures of central tendency (e.g., mean, median, mode) and variability (e.g., standard deviation, variance). 

Inferential statistics involve using sample data to make generalizations about a larger population, such as testing hypotheses or estimating population parameters.

Ques. What is the difference between a correlation and a regression analysis? (2 marks)

Ans. A correlation analysis examines the relationship between two variables, such as the strength and direction of the association. 

A regression analysis goes a step further by modelling the relationship between the variables, such as predicting one variable based on the other.

Ques. What is the difference between a discrete and a continuous variable? (2 marks)

Ans. A discrete variable can only take on specific, distinct values, such as the number of children in a family. 

A continuous variable can take on any value within a range, such as height or weight.

Ques. What is the difference between a Type I and a Type II error? (2 marks)

Ans. A Type I error occurs when a null hypothesis is rejected when it is actually true, meaning that the result is a false positive. 

A Type II error occurs when a null hypothesis is not rejected when it is actually false, meaning that the result is a false negative.

Ques. What are the assumptions underlying parametric statistics? (2 marks)

Ans. Parametric statistics assume that the data come from a normal distribution, have equal variances across groups, and are independent. 

These assumptions must be met in order for parametric tests, such as t-tests and ANOVA, to be valid.

Ques. The following are the heights of ten students in centimetres: 152, 168, 163, 167, 158, 172, 155, 162, 169, and 166. Calculate the mean, median, and mode. (3 marks)

Ans. The mean is (152 + 168 + 163 + 167 + 158 + 172 + 155 + 162 + 169 + 166) / 10 = 163.2 cm. 

The median is the middle value when the data are arranged in order, which is 165.5 cm.

The mode is the most frequent value, which is 163 cm.

Ques. The following are the scores of eight students on a test out of 100: 85, 92, 88, 70, 75, 90, 85, and 94. Calculate the interquartile range. (3 marks)

Ans. First, arrange the data in order: 70, 75, 85, 85, 88, 90, 92, 94. 

The first quartile is the median of the lower half of the data, which is (75 + 85) / 2 = 80. 

The third quartile is the median of the upper half of the data, which is (90 + 92) / 2 = 91. 

The interquartile range is the difference between the third and first quartiles, which is 91 - 80 = 11.

Ques. The following are the ages of ten employees in a company: 23, 28, 25, 30, 24, 29, 26, 31, 27, and 28. Calculate the coefficient of variation. (3 marks)

Ans. The coefficient of variation is the ratio of the standard deviation to the mean, expressed as a percentage. 

The mean is (23 + 28 + 25 + 30 + 24 + 29 + 26 + 31 + 27 + 28) / 10 = 27.1. 

The standard deviation is √[((23-27.1)2 + (28-27.1)+ (25-27.1)2 + (30-27.1)2 + (24-27.1)2 + (29-27.1)2 + (26-27.1)2 + (31-27.1)2 + (27-27.1)2 + (28-27.1)2) / 9] ≈ 2.3. 

Therefore, the coefficient of variation is (2.3 / 27.1) x 100% ≈ 8.5%.

Ques. The following are the daily sales figures in dollars for a store for five consecutive days: 1200, 950, 1450, 1000, and 900. Calculate the mean absolute deviation. (3 marks)

Ans. The mean is (1200 + 950 + 1450 + 1000 + 900) / 5 = 1100 dollars. 

The deviations from the mean are: 100, -150, 350, -100, and -200. 

Taking the absolute value of each deviation gives: 100, 150, 350, 100, and 200.

 The mean absolute deviation is the average of these absolute deviations, which is (100 + 150 + 350 + 100 + 200) / 5 = 180 dollars. 

Therefore, the mean absolute deviation from the mean daily sales is 180 dollars.

For Latest Updates on Upcoming Board Exams, Click Here: https://t.me/class_10_12_board_updates


Check-Out: 

Comments


No Comments To Show