BASES OF CLASSIFICATION
The bases or criteria for classifying data primarily depend on the
objectives and purpose of the inquiry. Generally, data can be classified
based on the following four bases:
1. Geographical: Area-wise or Regional.
2. Chronological: With respect to the occurrence of time.
3. Qualitative: With respect to some character or attribute.
4. Quantitative: With respect to numerical values or magnitudes.
Let’s discuss each of these in detail.
Geographical Classification
As the name suggests, in this classification, the basis is the geographical
or locational differences between various items in the data like States,
Cities, Regions, Zones, Areas, etc. For example, the yield of agricultural
output per hectare for different countries in a given period or the density
of the population (per square kilometer) in different cities of India.
Geographical classifications are usually presented either in an
alphabetical order (which is generally the case in reference tables) or
according to size or values to lay more emphasis on the important area or
region.
Chronological Classification
Chronological classification is one in which the data are classified based
on differences in time, e.g., the production of an industrial concern for
different periods; the profits of a big business house over different years;
the population of any country for different years. Time series data, which
are quite frequent in Economic and Business Statistics, are generally
classified chronologically, usually starting with the first period of
occurrence.
Qualitative Classification
When data are classified according to some qualitative phenomena which
are not capable of quantitative measurement like honesty, beauty,
employment, intelligence, occupation, sex, literacy, etc., the classification
is termed as qualitative or descriptive or with respect to attributes. In
qualitative classification, data are classified according to the presence or
absence of attributes in given units.
If data are classified into only two classes with respect to an attribute like
its presence or absence among various units, the classification is termed
as simple or dichotomous. Examples of such classification are classifying a
given population as honest or dishonest; male or female; employed or
unemployed; beautiful or not beautiful and so on.
However, if a given population is classified into more than two classes
with respect to a given attribute, it is said to be manifold classification. For
example, for the attribute intelligence, various classes may be genius,
very intelligent, average intelligent, below average and dull.
TYPES OF CLASS INTERVALS
Inclusive Classes
Both upper and lower class limits are included in the class, e.g. 40-
49.
Suitable for discrete variables like test scores, accidents, etc. taking
only integral values.
Doesn't work for continuous variables like height, weight, etc.
Exclusive Classes
Upper limit excluded from the class, included in next class, e.g. 20-
25 means 20 ≤ X < 25.
Needed for continuous variables to avoid gaps.
Can record age on late or next birthday to make discrete then
convert to exclusive.
Open End Classes
One limit (lower or upper) not specified, e.g. age above 60.
Avoid if possible as mid-value unclear, issues for analysis and
graphs.
May be needed for economic/medical data with extreme values.
Estimate mid-value based on previous/next class.
Key Points
Nature of variable guides inclusive vs exclusive classes.
Exclusive classes needed for continuity, analysis of continuous data.
Class limits should give even distribution around mid-point.
Avoid open end classes if possible.
Proper choice of class type and limits is critical to reveal distribution
characteristics accurately and allow valid analysis.
BIVARIATE FREQUENCY DISTRIBUTION
Simultaneous classification on two variables for the same population
Results in a two-way table called bivariate frequency table
Variables X and Y grouped into m and n classes respectively
Gives m x n cells with frequency f(x,y) for each (x,y) pair
Marginal Distributions
Frequency distribution of X values and totals fx is marginal
distribution of X
Frequency distribution of Y values and totals fy is marginal
distribution of Y
Conditional Distributions
Distribution of X for a fixed value of Y is conditional distribution of X
Distribution of Y for a fixed value of X is conditional distribution of Y
Key Points
Bivariate table shows visual relationship between two variables
Marginal distributions give individual distributions of X and Y
Conditional distributions fix one variable and show distribution of
the other
Quantitative correlation measured by correlation coefficient
Bivariate frequency analysis allows joint study of two related variables and
examination of their marginal and conditional distributions.
WHAT IS TABULATION
Tabulation is the systematic presentation of data in rows and columns to
summarize and organize information. It is an intermediate step between
data collection and analysis.
Parts of a Table
Table number for identification
Title describing contents concisely
Head notes giving units of measurement
Column captions and row stubs with headings
Body with data values, totals, grand totals
Footnotes for extra details
Source note indicating data origin
Requisites of a Good Table
Simple, compact, complete, and self-explanatory
Logical classification showing relationships
Accurate and error-free
Attractive layout aiding comprehension
Data in useful orders like magnitude, chronological
Adequate totals, ratios, percentages for interpretation
Types of Tables
General reference tables for record keeping
Special summary tables for analytical comparisons
Primary tables with original data values
Derived tables with ratios, percentages, etc.
Simple one-way tables of one characteristic
Complex multivariate tables
Tabulation summarizes data efficiently, revealing relationships and trends.
Judgment and experience guide table preparation to achieve relevance,
completeness, accuracy andvisual appeal.