
Content Writer
Categorical data in statistics refers to the data which is categorized according to its categorical variables. It is a countable term that is obtained from qualitative data analysis. Grouped data is an example of categorical data.
- It is possible to deduce categorical data by grouping them into suitable intervals.
- Probability tables provide a summary of these data.
- The term "categorical data" often refers to datasets during data analysis.
- It is also referred to as an enumerated data.
- Each possible value of categorical data is referred to as level.
- Data is a set of facts and figures in any form.
- The most common example includes rolling a dice.
- Rolling a dice has six outcomes, including 1, 2, 3, 4, 5, and 6, which are all categorical data.
Read More: Universal Set
Key Terms: Numerical data, Categorical Data, Qualitative Data, Categorical Variables, Nominal Data, Ordinal Data, Probability, Observations
Categorical Data
[Click Here for Sample Questions]
Categorical data is a type of data comprising a categorical characteristic such as a person's gender, hometown, etc. Examples of categorical data include travel method to school, favourite sport, school postcode, birthdate, and many more.
- The birthdate and postcode in the example above both contain a number system.
- In this method, information is identified on the basis of labels or names.
- Categorical data are grouped into categories.
- The data represented by two values is called a dichotomous variable or binary variable.
- The type of data represented by more than two values is called polytomous variables.
Read More: LCM of Two Numbers
Types of Categorical Data
[Click Here for Sample Questions]
Data that can be categorized or grouped under a category are called categorical data. Bar graphs and pie charts are the most effective ways to show categorical data. There are two kinds of categorical data, which are as follows:
Nominal Data
Numerical data does not provide a numerical value to the variables, so this type of information is called nominal data. A nominal scale is another term for it.
- There is no way to order or measure nominal data.
- Nevertheless, nominal data can sometimes be both qualitative and quantitative.
- Set Symbols, words, letters, and gender are some examples of nominal data.
Ordinal Data
An ordinal dataset is a dataset organized in accordance with its natural order. Ordinal data differs from nominal data because it can't determine if the two differ. The most common way of presenting it is through a bar chart.

Categorical Data
Example of Categorical Data
[Click Here for Sample Questions]
Categorical data is shown in two-way tables by counting the number of observations that fall into two categories for two variables, one arranged in rows and the other in columns. If, for example, 20 individuals were asked to identify the colour of their hair and eyes during a survey, how would they do? It would look something like this:
| Eye Colour | |||||
|---|---|---|---|---|---|
| Hair Colour | Blue | Green | Brown | Black | Total |
| Blonde | 2 | 1 | 2 | 1 | 6 |
| Red | 1 | 1 | 2 | 0 | 4 |
| Brown | 1 | 0 | 4 | 2 | 7 |
| Black | 1 | 0 | 2 | 0 | 3 |
| Total | 5 | 2 | 10 | 3 | 20 |
According to marginal distributions, the number of individuals in each row or column is determined without considering how other linear equations in two variables influence the number of individuals with blue eyes.
- In analysing two-way tables, percentages are often used as opposed to simple counts because simple counts are often difficult to analyse.
- Here are four examples of individuals with red hair.
- The 20 observations mean that 20% of the people who were surveyed were redheaded.
Read More: Place Values
Categorical Variables
[Click Here for Sample Questions]
Variables with a fixed number of possible values are put into the category of categorical variables in statistics. These variables typically take name or label values. Some example of categorical variables are as follows:
- Wall colours include red, blue, pink, green, etc.
- The gender of a person, such as male, female, or transgender.
- A, B, O, AB, etc., are the blood groups of a person.
Categorical distributions are probability distributions correlated with categorical variables.
Read More: Independent Events in Probability
Things to Remember
[Click Here for Sample Questions]
- Categorical data represents a type of data that can be categorised into groups.
- There are categorical variables such as race, gender, age, and educational level.
- It is often more informative to group the variables into a relatively small number of groups.
- Categorical data is qualitative. Instead of numbers, words are used to describe events.
- To analyse categorical data, modes and medians are used.
- Nominal data is analysed in terms of modes, and ordinal data is analysed using both.
- It is possible to express categorical data numerically (e.g., "1" indicates Yes and "2" indicates No), but these numbers do not have mathematical significance.
Read More: Frequency Distribution Table Statistics
Sample Questions
Ques: Find the mean deviation about the median for the following data:
(5 Marks)
| xi | 3 | 6 | 9 | 12 | 13 | 15 | 21 | 22 |
|---|---|---|---|---|---|---|---|---|
| fi | 3 | 4 | 5 | 2 | 4 | 5 | 4 | 3 |
| c.f. | 3 | 7 | 12 | 14 | 18 | 23 | 27 | 30 |
Now, N=30 which is even.
In this case, the median refers to the mean of the 15th and 16th observations. For each of these observations, the corresponding observation is 13 . Both of these observations fall within the cumulative frequency 18.
Therefore, Median M=(15th observation +16th observation) /2=(13+13)/2=13
Here are the absolute deviations from the median, i.e., |xi-M| are shown in the table.
| xi-M | 10 | 7 | 4 | 1 | 0 | 2 | 8 | 9 |
|---|---|---|---|---|---|---|---|---|
| fi | 3 | 4 | 5 | 2 | 4 | 5 | 4 | 3 |
| f|xi-M| | 30 | 28 | 20 | 2 | 0 | 10 | 32 | 27 |
We have ∑i=\(\displaystyle\sum_{i=1}^{8}\)fi=30 and ∑i=\(\displaystyle\sum_{i=1}^{8}\)fi|xi-M|=149
Therefore
- D. (M) =1 /N\(\displaystyle\sum_{i=1}^{8}\)fi|xi-M| =1/30×149=4.97
Ques: Find the mean deviation about the mean for the following data.
(5 Marks)
Ans: We make the following from the given data:
| Marks obtained | Number of students | Mid-points | fixi | xi-x‾ | fixi-x‾ |
|---|---|---|---|---|---|
| 10-20 | fi | xi | |||
| 20-30 | 2 | 15 | 30 | 30 | 60 |
| 30-40 | 8 | 25 | 75 | 20 | 60 |
| 40-50 | 14 | 45 | 280 | 10 | 80 |
| 50-60 | 8 | 55 | 440 | 10 | 80 |
| 60-70 | 3 | 65 | 195 | 20 | 60 |
| 70-80 | 2 | 75 | 150 | 30 | 60 |
| 40 | 1800 | 400 |
Here, N=i=\(\displaystyle\sum_{i=1}^{7}\)fi=40,i=\(\displaystyle\sum_{i=1}^{7}\)fi|xi|=1800,i=\(\displaystyle\sum_{i=1}^{7}\)fi|xi-x‾|=400
Therefore, x‾=1 /N\(\displaystyle\sum_{i=1}^{7}\)fi|xi|=1800/40=45
And, M.D. (x‾)=1 /N\(\displaystyle\sum_{i=1}^{7}\)fi|xi-x‾|=1/40×400=10
Ques: Find the mean deviation about the mean for the following data: (4 Marks)
6,7,10,12,13,4,8,12
Ans: We will proceed step-wise and get the following:
Step 1: Mean of the given data is:
x‾=(6+7+10+12+13+4+8+12)/8=72/8=9
Step 2: Deviations from the mean of the respective observations x‾, i.e., |xi-x‾| are 6-9,7-9,10-9,12-9,13-9,4-9,8-9,12-9, or -3,-2,1,3,4,-5,-1,3
Step 3: Here are the absolute deviations from the median, i.e., |xi-x‾| are3,2,1,3,4,5,1,3
Step 4: The required variation from the mean is
M.D. (x‾) =\(\displaystyle\sum_{i=1}^{8}\)|xi-x‾|/8 =(3+2+1+3+4+5+1+3)/8=22/8=2.75
Ques: Calculate the mean deviation about median for the following data:
(5 Marks)
Ans: Form the following from the given data:
| Class | Frequency | Cumulative frequency | Mid-points | ?xi- Med. | fi?xi- Med. ? |
|---|---|---|---|---|---|
| 0-10 | fi | (c.f.) | xi | ||
| 10-20 | 7 | 6 | 5 | 23 | 138 |
| 20-30 | 15 | 13 | 15 | 13 | 91 |
| 30-40 | 16 | 28 | 25 | 3 | 45 |
| 40-50 | 4 | 44 | 35 | 7 | 112 |
| 50-60 | 2 | 50 | 55 | 27 | 58 |
The class interval containing Nth 2 or 25th item is 20-30. Therefore, 20-30 is the median class. We know that
Median =l+((N/2-C)/f)xh Here l=20,C=13,f=15,h=10 and N=50
Therefore,Median =20+(25-13)/15×10=20+8=28
Thus, Mean deviation about median is given by:
M.D. (M) =1 /N\(\displaystyle\sum_{i=1}^{6}\)fi|xi-M|=1/50×508=10.16
Ques: Find mean deviation about the mean for the following data : (3 Marks)
xi 2 5 6 8 10 12 fi 2 8 10 7 8 5
Ans: The given data and append other columns after calculations:
| xi | fi | fixi | xi-x‾ | fixi-x‾ |
|---|---|---|---|---|
| 2 | 2 | 4 | 5.5 | 11 |
| 5 | 8 | 40 | 2.5 | 20 |
| 6 | 10 | 60 | 1.5 | 15 |
| 8 | 7 | 56 | 0.5 | 3.5 |
| 10 | 8 | 80 | 2.5 | 20 |
| 12 | 5 | 60 | 4.5 | 22.5 |
| 40 | 300 | 92 |
N=i=\(\displaystyle\sum_{i=1}^{6}\)fi=40, i=\(\displaystyle\sum_{i=1}^{6}\)fixi=300, i=\(\displaystyle\sum_{i=1}^{6}\)fi|xi-x‾|=92
Therefore x‾=1 /N\(\displaystyle\sum_{i=1}^{6}\)fixi=1/40×300=7.5
and M. D. (x‾)=1 /N\(\displaystyle\sum_{i=1}^{6}\)fi|xi-x‾|=1/40×92=2.3
Ques: Find the mean deviation about the mean for the following data: (4 Marks)
12,3,18,17,4,9,17,19,20,15,8,17,2,3,16,11,3,1,0,5
Ans: We have to first find the mean (x‾) of the given data
x‾=1/20\(\displaystyle\sum_{i=1}^{20}\)xi=200/20=10
The absolute values of the deviations from mean, i.e., |xi-x‾| are 2,7,8,7,6,1,7,9,10,5,2,7,8,7,6,1,7,9,10,5
Therefore ∑i=\(\displaystyle\sum_{i=1}^{20}\)|xi-x‾|=124 and M.D. (x‾)=124/20=6.2
Ques: Find the mean deviation about the median for the following data: (3 Marks)
3,9,5,3,12,10,18,4,7,19,21.
Ans: Here the number of observations is 11 which is odd. Arranging the data into ascending order, we have 3,3,4,5,7,9,10,12,18,19,21
Now Median =(11+1)th/2 or 6th observation =9
The values of the respective deviations from the median, i.e., |xi-M| are 6,6,5,4,2,0,1,3,9,10,12
Therefore, i=\(\displaystyle\sum_{i=1}^{11}\)|xi-M|=58
and, M.D. (M)=1/11\(\displaystyle\sum_{i=1}^{11}\)|xi-M|=1/11×58=5.27
Ques. What is the difference between nominal categorical data and ordinal categorical data. [3 marks]
Ans. The difference between nominal categorical data and ordinal categorical data are as follows:
| Nominal Categorical Data | Ordinal Categorical Data |
|---|---|
| Nominal Categorical Data are type of data that are non-parametric and non-ordered data. | Ordinal Categorical Data are type of data that are non-parametric ordered data. |
| It is used to group similar objects under a similar category. | It is based on the opinion of people. |
| For eg: gender and hair colour | For eg: people point of view in a survey |
Ques. Let’s say you are having a party and want to make sure everyone has coffee to drink. So you send out a survey asking people what their favorite coffee is, and you put the answers into a table like the one below: [1 mark]
| Favorite Coffee | Frequency |
|---|---|
| Latte | 80 |
| Espresso | 76 |
| Cappuccino | 12 |
| Black Coffee | 10 |
Is the data in the table categorical?
Ans. Yes. It is categorical data because it is broken up into groups, like favorite coffee. Each name of the type of coffee is referred to as level.
Ques. If the ratio of the mode to the median is 4 : 3 for a dataset. Find the ratio of mode to mean. [2 marks]
Ans. The empirical relationship between mean mode and median is
Mode = 3 Median – 2 Mean.
- Let the modal value be 4x, and the median be 3x, then
- 4x = 3.3x – 2 Mean
- 2 Mean = 9x – 4x
- Mean = 5x / 2
- Mode : Mean = 4x / (5x/2) = 8 / 5
- The ratio of mode to mean is 8 / 5.
Ques. Calculate the mean from the data showing marks of students in a class in a test: 60, 50, 55, 70, 80. [2 marks]
Ans. Given marks: 60, 50, 55, 70, 80
- Here, the number of data values = 5
- We know that:
- Mean = Sum of data values/Total number of data values
- (60 + 50 + 55 + 79 + 80) / 5
- 315 / 5
- 63
Therefore, the mean for the given data is 63.
Also Check:






Comments