
Education Journalist | Study Abroad Lead
A sample is a cluster or group of elements chosen from a population. It helps to estimate what the entire population is doing, without having to survey everyone. Hence, it is a method of saving plenty of time and money. The attributes used to define a population are referred to as the parameters and the properties of a sample data are known as statistics. The concept of both the population and sample are essential components of statistics. A sample statistic is a portion of the data collected from a fraction of a population. It should define the population as a whole and not reflect any inclination toward a specific attribute. The concept of samples is being used for studies and measurements by scientists, marketers, government agencies, economists, research groups, etc. It involves important topics such as hypothesized mean, mean standard deviation, distribution of means, and various other concepts.
| Table of Content |
Key takeaway: Sample, mean, population, attributes, statistics, cluster, hypothesized mean, standard deviation, distribution, unbiased, observations, median, frequency, sampling
What is sample in statistics?
[Click Here for Sample Questions]
A sample is generally a smaller and manageable version of a larger group. It is a subset having the attributes of a larger population. A population, in general, is the total number of observations contained in a given group or context. When the sizes of the population are huge to include all possible members or observations, then the concept of sampling is used in statistical testing.
The image below shows the relation between a population and the sample.

Sample of a Population
As shown in the above figure, a sample is an unbiased number of observations carried out from a population. In other words, it is a portion or fraction of the whole group and functions as the subset of a population.
Sample mean
[Click Here for Sample Questions]
In statistics, the mean is termed as one of the measures of central tendency. It is similar to the average of a given set of values. It represents the equal value distribution for the data set provided. It is another term for “median.” Generally, the sample mean is an average sample size and is only a tiny component of the total population. It is helpful as it helps to evaluate what is being done by the entire population without conducting any survey.
Sample Mean formula
[Click Here for Sample Questions]
The sample mean of any population can be determined by adding all the observations and then dividing the resultant by the total number of observations made i.e. ‘n’. The formula for evaluating the sample mean of a population is given as:
\(\bar{X}\) = (\(\sum\) xi) / n or \(\bar{X}\) = 1/ n * (\(\sum\)xi)
From the above equation, it can be concluded that the concept of average and sample mean is nearly the same. However, the only difference existing between both of them is their symbols.
In the above equation,
\(\bar{X}\) = stands for ‘sample mean’
\(\sum\)= summation notation
xi n = all of the values of x
n = number of items in sample mean
Hypothesized mean
[Click Here for Sample Questions]
In statistics, hypothesis testing is a technique in which an analyst tests an assumption or a hypothesis about a population parameter. It is used to evaluate the probability of a hypothesis by using sample data that may have come from a larger population or any data-generating process.
Its process involves the use of a random population sample for testing two different hypotheses which are:
- Null hypothesis (H0)
- Alternative hypothesis (Hi)
The null hypothesis is the hypothesis of equality between population parameters whereas the alternative hypothesis is the opposite of a null hypothesis.
Estimating the mean
For calculating the mean of a given population or group, follow the steps listed below:
Step 1: First create a new column into the table and write the middle value (midpoint) of each group.
Step 2: Now multiply middle values by the frequency of that group and Step 3: Create a new column for (midpoint x frequency) values and add the results obtained from the above step into the column.
Step 4: At last, divide that value obtained by the net frequency to obtain the mean.
Sampling distribution
[Click Here for Sample Questions]
A sampling distribution also referred to as a finite-sample distribution, is similar to the probability distribution of a statistic that is selected from random samples of a population. It illustrates the distribution of frequencies for how to spread apart various outcomes for a specific population.
It relies on numerous aspects like statistics, sample size, sampling process, the overall population, etc. and helps in calculating the statistics such as means, ranges, variances, and standard deviations for the sample provided.
The variability of a sampling distribution is determined by the number of observations in a population and a sample, and the approach used to draw the sample sets.
Solved examples
[Click Here for Sample Questions]
Example 1: An exam of 100 marks was conducted in Class 10th of 30 students and marks obtained by them have been recorded in the table below.
| Marks obtained | No. of students |
|---|---|
| 10 | 1 |
| 20 | 1 |
| 36 | 3 |
| 40 | 4 |
| 50 | 3 |
| 56 | 2 |
| 60 | 4 |
| 70 | 4 |
| 72 | 1 |
| 80 | 1 |
| 88 | 2 |
| 92 | 3 |
| 95 | 1 |
You have to calculate the mean of the obtained marks by the students.
Solution: We will take the marks obtained by the students as xi and number of students as fi. We will then calculate the product of fixi for each observation individually.
| Marks obtained (xi) | No. of students (fi) | fixi |
|---|---|---|
| 10 | 1 | 10 |
| 20 | 1 | 20 |
| 36 | 3 | 108 |
| 40 | 4 | 160 |
| 50 | 3 | 150 |
| 56 | 2 | 112 |
| 60 | 4 | 240 |
| 70 | 4 | 280 |
| 72 | 1 | 72 |
| 80 | 1 | 80 |
| 88 | 2 | 176 |
| 92 | 3 | 276 |
| 95 | 1 | 95 |
| Total | 30 | 1779 |
Using the formula for mean we will now calculate the mean of marks obtained by the students:
Mean,\(\bar{X}\) = \(\sum f_i. x_i/ \sum f_i\)
=> 1779/30
=> 59.3
Hence the mean marks obtained is 59.3
Example: A survey was conducted in a locality of 25 households for daily expenditure on food. The table below contains the details of the survey. You have to evaluate the daily expenditure on food by an appropriate method.
| Daily expenditure | No. of household |
|---|---|
| 100-150 | 4 |
| 150-200 | 5 |
| 200-250 | 12 |
| 250-300 | 2 |
| 300-350 | 2 |
Solution: From the given question, we can take the value of assumed mean a, and class interval h as 225 and 50 respectively.
| Daily expenditure (in Rs.) | No. of households (fi) | Class mark (xi) | xi – 225 | ui = (xi – 225) / 50 | fiui |
|---|---|---|---|---|---|
| 100-150 | 4 | 125 | -100 | -2 | -8 |
| 150-200 | 5 | 175 | -50 | -1 | -5 |
| 200-250 | 12 | 225 | 0 | 0 | 0 |
| 250-300 | 2 | 275 | 50 | 1 | 2 |
| 300-350 | 2 | 325 | 100 | 2 | 4 |
| Total | ∑fi = 25 | ∑fiui = -7 |
Mean, \(\bar{X}\) = a + h * (\(\sum f_i u_i* \sum f_i\))
Putting the values of a and h in the above equation,
\(\bar{X}\)= 225 + 50 * (-7/25)
=> 225 -14
=> 211
Hence, main daily expenditure on food calculate will be Rs. 211
Things to remember
- The sample mean is referred to as the average value found in a sample. It is simply a small portion of the whole population. In other words, mean refers to “average.”
- Sample mean is determined using the formula \(\bar{X}\) = 1/ n * (Σ xi) where n is the total number of observations.
- Hypothesis testing is an important approach in statistics used to evaluate two mutually exclusive statements about a population that specify which statement is most suitable and supported by the sample data.
- A sampling distribution is typically a graph of a statistic for the sample data. It’s the probability distribution of a statistic achieved from a larger number of samples taken from a specific population. For a given population, the sampling distribution is the allocation of frequencies of a range of distinct outcomes that could occur for a statistic of a population.
Sample Questions
Ques: Elaborate on how does the sampling distribution work? (2 marks)
Ans: The working of the sampling distribution has been explained below:
Step 1: From the population provided, pick a random sample of a specific size.
Step 2: Now, calculate a statistic for the sample. For instance the mean, median, or standard deviation.
Step 3: Create a frequency distribution for each sample statistic obtained in the above step.
Step 4: At last, plot the frequency distribution of each sample statistic obtained in the above step.
The resultant graph obtained will be the sampling distribution.
Ques: The following data is collected by a group of students regarding the number of plants in 20 houses in a locality. Calculate the mean number of plants per house. (3 marks)

Ans: Since the values of xi and fi are very small hence we can use direct method for calculating the mean.
| No. of plants | Class mark (xi) | No. of houses (fi) | fi.xi |
|---|---|---|---|
| 0-2 | 1 | 1 | 01 |
| 2-4 | 3 | 2 | 06 |
| 4-6 | 5 | 1 | 05 |
| 6-8 | 7 | 5 | 35 |
| 8-10 | 9 | 6 | 54 |
| 10-12 | 11 | 2 | 22 |
| 12-14 | 13 | 3 | 39 |
| Total | ∑fi = 20 | ∑fi.xi = 20 | |
We know,
Mean, \(\bar{X}\) =\(\sum f_i. x_i / \sum f_i\)
= 162 / 20
=> 8.1
Hence, there are a mean of 8.1 plants per house
Ques: The following table contains the distribution of daily wages of 50 labors of a construction company. Determine the mean daily wages of the labors. (3 marks)

Ans: Since the data is large thus we can use the step-deviation approach for calculating the mean.
In the given question, a=150 and h=20
| Class interval | Frequency (fi) | Class marks | ui = (xi – a)/h | Fiui |
|---|---|---|---|---|
| 100-120 | 12 | 110 | -2 | -24 |
| 120-140 | 14 | 130 | -1 | -14 |
| 140-160 | 8 | 150 | 0 | 0 |
| 160-180 | 6 | 170 | 1 | 6 |
| 180-200 | 210 | 190 | 2 | 20 |
| ∑fi = 50 | ∑fiui = -12 |
Using the formula mean for step-deviation,
\(\bar{X}\) = a + h (\(\sum f_i u_i/ \sum f_i\))
= 150 + 20 (-12/50)
= 150 - (24/5)
= (750-24) / 5
= 726/5
=> 145.20
Therefore, the mean daily wages of the labors is Rs. 145.20
Ques: The table below contains the distribution values of the daily pocket allowance of children of a society. Given the mean pocket allowance as Rs. 18, calculate the missing frequency f. (5 marks)

Ans:
| Daily pocket allowance (in Rs.) | Class mark (xi) | No. of children (fi) | di = xi - 18 | fidi |
|---|---|---|---|---|
| 11-13 | 12 | 7 | -6 | -42 |
| 13-15 | 14 | 6 | -4 | -24 |
| 15-17 | 16 | 9 | -2 | -18 |
| 17-19 | 18 | 13 | 0 | 0 |
| 19-21 | 20 | F | 2 | 2f |
| 21-23 | 22 | 5 | 4 | 20 |
| 23-25 | 24 | 4 | 6 | 24 |
| Total | ∑fi = 44 + f | ∑fidi = 2f - 40 |
Using the mean formula, \(\bar{X}\) = a + (\(\sum f_i d_i / \sum f_i\))
=> 18 = 18 + [(2f – 40) / (44 + f)]
=> 0 = (2f – 40) / (44 + f)
=> 2f – 40 = 0
=> 2f = 40
=> f = 40/2
=> 20
Therefore, the value of missing frequency f is 20.
Ques: The total marks obtained by few students in mathematics exam are 101, 161, 155, 96 and 83. Evaluate the sample mean marks? (2 marks)
Ans: The formula for mean is,
\(\bar{X}\) = \(\frac{\sum_{i=1}^n x_i}{n}\)
Form the given question, the total number of terms, n = 5
Therefore,
Mean, \(\bar{X}\) = (101+161+155+96+83) / 5
= 119.2
Hence the sample mean marks obtained by the students is 119.2
Ques: Find the sample mean for the following set of numbers: 12, 13, 14, 16, 17, 40, 43, 55, 56, 67, 78, 78, 79, 80, 81, 90, 99, 101, 102, 304, 306, 400, 401, 403, 404, and 405. (3 marks)
Ans: Adding all the given set of numbers we will obtain,
=> 12 + 13 + 14 + 16 + 17 + 40 + 43 + 55 + 56 + 67 + 78 + 78 + 79 + 80 + 81 + 90 + 99 + 101 + 102 + 304 + 306 + 400 + 401 + 403 + 404 + 405
=> 3744
Since the data set have 26 items hence the value of n will 26 i.e. n=26
For calculating the mean, divide the sum of all the numbers obtained previously by the total number of items:
=> 3744/26
=> 144
Therefore, the sample mean of the given set of numbers is 144.
Ques: Calculate the mean and variance of the following frequency distribution: (3 marks)

Ans: Let’s take the assumed mean, A = 25.5
In the given question, value of h = 10
| Classes | xi | yi = (xi – 25.5)/10 | fi` | fi yi | fi yi2 |
|---|---|---|---|---|---|
| 1-10 | 5.5 | -2 | 11 | -22 | 44 |
| 10-20 | 15.5 | -1 | 29 | -29 | 29 |
| 20-30 | 25.5 | 0 | 18 | 0 | 0 |
| 30-40 | 35.5 | 1 | 4 | 4 | 4 |
| 40-50 | 45.5 | 2 | 5 | 10 | 20 |
| 50-60 | 55.5 | 3 | 3 | 9 | 27 |
| 70 | -28 | 124 | |||
x′ = (fi yi) / yi = -28/70 = -0.4
Using the formula for mean, \(\bar{X}\)
\(\bar{X}\) = \(\frac{\sum_{i=1}^n x_i}{n}\)
= 25.5 + (-10) (10.4)
=> 21.5
Using the formula for variance, (σ2)
σ2 =
\(\frac{h}{N} \sqrt{Nf_iy_i^2 - (f_iy_i)^2}^2\)= [(10*10) / (70*70)] [70(124) - (-28)2]
= 70(124) / (7*7) – (28*28) / (7*7)
= (1240 / 7) – 16
=> 161
Hence the mean and variance of the given distribution is 21.5 and 161 respectively.
Ques: There are two factories A and B, which are producing bulbs. The table below has the values of life of bulbs produced by the factories. From the point of view of length of life, determine which factory’s bulbs are more consistent. (5 marks)

Ans: From the given question we can take the value of h=100, and the assumed mean, A=800
| Length of life (in hrs) | Mid values (xi) | yi = (xi – A)/10 | Factory A | Factory B | ||||
|---|---|---|---|---|---|---|---|---|
| fi | fi yi | fi yi2 | fi | fi yi | fi yi2 | |||
| 550-650 | 600 | -2 | 10 | -20 | 40 | 8 | -16 | 32 |
| 650-750 | 700 | -1 | 22 | -22 | 22 | 60 | -60 | 60 |
| 750-850 | 800 | 0 | 52 | 0 | 0 | 24 | 0 | 0 |
| 850-950 | 900 | 1 | 20 | 20 | 20 | 16 | 16 | 16 |
| 950-1050 | 1000 | 2 | 16 | 32 | 64 | 12 | 24 | 48 |
| Total | 120 | 10 | 146 | 120 | -36 | 156 | ||
Using the mean formula
For factory A
Mean, \(\bar{X}\) = 800 + (10 / 120) * 100
⇒ 816.67 hrs
Standard deviation, S.D = (100/120) √[(120*146)-100]
⇒ 109.98
Hence, coefficient of variation = (S.D / \(\bar{X}\))*100
= (109.98 / 816.67)* 100
⇒ 13.47
Similarly we can calculate the values for factory B
For factory B
Mean,\(\bar{X}\) = 800 + (-36 / 120)100
⇒ 700
Standard deviation, S.D = (100/120) √[120(156)-(-36)2]
⇒ 110
Coefficient of variation = (S.D / \(\bar{X}\))*100
= (110/770)100
⇒ 14.29
As the Coefficient of variation of factory B is greater than that of factory A. This means that Factory B has more variability and thus the bulbs of factory A are more consistent than B.
Ques: Calculate the Standard error of the following dataset of heights given: (5 marks)

Ans: Mean of the given dataset = (170.5 + 161 + 160 + 170 + 150.5) / 5
=> 162.4
Calculating the deviation from the mean calculated above:
170.5 – 162.4 = -8.1
161 – 162.4 = 1.4
160 – 162.4 = 2.4
170 – 162.4 = -7.6
150.5 – 162.4 = 11.9
Squaring and adding the above values individually:
= (-8.1)2 + (1.4)2 + (2.4)2 + (-7.6)2+ (11.9)2
= 65.61 + 1.96 + 5.76 + 57.76 + 141.61
=> 272.7
Dividing the above result by the sample size -1 i.e. n-1=4 as the sample contains five items
= (272.7) / 4
=> 68.175
Standard deviation, S.D. = \(\sqrt (\)68.175)
=> 8.257
Dividing the standard deviation calculated by the square root of the sample size:
Standard error, S.E = S.D / \(\sqrt(n)\)
=> 8.257 / \(\sqrt(5)\)
=> (8.257) / (2.236)
=> 3.693
Hence, the standard error for the given dataset of heights is 3.693.
Ques: The table below contains the literacy rate of 35 cities in percentage. Calculated the mean literacy rate for the following: (3 marks)

Ans: From the given table we can take the values of a and h i.e. a=70 and h=10
| Literacy rate (in %) | Class marks (xi) | No. of cities (fi) | ui = (xi -70) / 10 | fiui |
|---|---|---|---|---|
| 45-55 | 50 | 3 | -2 | -6 |
| 55-65 | 60 | 10 | -1 | -10 |
| 65-75 | 70 | 11 | 0 | 0 |
| 75-85 | 80 | 8 | 1 | 8 |
| 85-95 | 90 | 3 | 2 | 6 |
| Total | ∑fi = 35 | ∑fiui = -2 |
Using mean formula,
\(\bar{X}\) = a + (\(\sum f_i u_i / \sum f_i\)) * h
=> 70 + [(-2 *10) / 35]
=> 70 – 0.57
=> 69.43%
Hence, the mean literacy rate is 69.43%
Related articles







Comments