DoAssignment.ca
1.5 · Organize and graph quantitative data
Learn to organize and graph quantitative data through clear examples and targeted practice.
Athabasca University MATH 215: Introduction to Statistics
Descriptive Statistics
A practical guide to turning numerical observations into clear displays
A list of numbers can be accurate but difficult to understand. Organizing quantitative data means arranging numerical observations so patterns are easier to see. A graph can show where values cluster, how much they vary, whether the distribution is balanced or uneven, and whether any values stand apart. This lesson focuses on choosing and making simple displays that suit numerical data. A graph summarizes the observations; it does not explain why the values occurred.
What you will learn
- Distinguish quantitative data from categories and identify the observational units and variable.
- Organize quantitative observations with a frequency distribution.
- Construct and interpret a histogram, dot plot, or stem-and-leaf display.
- Describe a distribution using its shape, centre, spread, and any unusual values.
1. Identify the data and choose a display
An observational unit is the person or object on which a measurement is made. A variable is the characteristic recorded for each unit. Quantitative data are measurements or counts for which arithmetic makes sense, such as commute time in minutes or number of messages received. By contrast, a category such as a person's preferred route is not a numerical measurement, even if categories are assigned number codes.
A sample is the group of units actually observed. The individual recorded numbers are observations, and the collection of those numbers is a data set. Before graphing, check that the values refer to the same variable, use consistent units, and have been recorded correctly. Decide whether the values are counts, which take separate whole-number values, or measurements, which may take many values along a scale.
A dot plot places a dot above a number line for each observation. It is especially useful for a small or moderate data set and keeps individual values visible. A stem-and-leaf display also preserves the original values while grouping them by place value. A histogram groups numerical values into intervals called classes or bins; it is useful when a data set is too large for a dot plot to remain clear.
For any graph, choose a scale that covers all observations, label the variable and units, and make the display readable. Do not choose intervals that hide important features or change them partway through the graph. For a histogram, classes should cover the data range without gaps or overlap, so each observation belongs to exactly one class.
- Identify the units, variable, and units of measurement before drawing.
- Use a dot plot or stem-and-leaf display when seeing individual observations is useful.
- Use a histogram to summarize a larger set of numerical observations.
2. Build a frequency distribution
A frequency distribution lists classes and the number of observations in each class. Frequency means count. For example, if five observations fall in a particular interval, that interval's frequency is five. The class intervals must be defined precisely: a value on a boundary must not fit into two classes. One clear convention is to include the lower endpoint and exclude the upper endpoint, except that the final class can include its upper endpoint.
To choose classes, first find the smallest and largest observations and note the overall range, meaning largest minus smallest. Select a manageable number of equal-width classes that cover this range. There is no single class count that suits every data set. Too few classes can conceal structure; too many can make the graph almost as cluttered as the raw list. Write the class boundaries or an unambiguous endpoint rule before tallying.
Tally each observation into one class, then count the tallies. Check that the class frequencies add to the total number of observations, usually written as . This is an important error check: every recorded observation must be counted once, not skipped or counted twice. A relative frequency is the share of observations in a class. It can be written as a decimal or multiplied by to give a percentage.
In a frequency histogram, class intervals appear along the horizontal axis and frequencies on the vertical axis. Draw a bar for each class; adjacent bars touch because the numerical intervals are continuous across their boundaries. The bar height represents the frequency. When class widths are equal, comparing heights directly compares counts. If widths differ, bar area rather than height should represent frequency, so equal-width classes are the simpler choice for an introductory hand-drawn histogram.
- Use a consistent rule for class endpoints.
- Check that frequencies sum to the sample size.
- Label both axes, including the frequency or relative-frequency scale.
3. Read the overall pattern
A distribution is the pattern of values in a data set. Describe a graph by considering its shape, centre, spread, and unusual features. Shape includes whether values form one main cluster or several, and whether the pattern is roughly balanced or has a longer tail to one side. A longer tail toward larger values is called right-skewed; a longer tail toward smaller values is called left-skewed. Use these descriptions only when the graph supports them.
The centre is the general middle or typical location of the observations. Spread describes how much the values vary. At this stage, these can be described in the units of the variable and by referring to the graph; a formal numerical summary is not required to make a useful first description. An isolated value far from the others may be an unusual observation. Check such a value against the original record before deciding it is real; a graph alone cannot establish that it is an error.
A graph is sensitive to choices such as class width, starting point, and axis scale. A different reasonable grouping can make a pattern look somewhat different. Therefore, use clear scales, state the class intervals, and avoid making a strong claim from a small visual difference. Graphs summarize the data that were collected, not every possible value or a reason for the pattern.
- Describe shape, centre, spread, and possible unusual values in context.
- Do not call a graph balanced or skewed without a visible pattern to support it.
- A display summarizes observed data; it does not establish cause and effect.
Commute-time frequency distribution
| Commute time (minutes) | Frequency | Relative frequency |
|---|---|---|
| 10 to under 20 | 2 | 25.0% |
| 20 to under 30 | 4 | 50.0% |
| 30 to under 40 | 1 | 12.5% |
| 40 to 50 | 1 | 12.5% |
| Total | 8 | 100.0% |
Worked example
From commute times to a histogram
Eight students report one-way commute times, in minutes: 12, 18, 21, 23, 24, 27, 31, and 44. Organize the data into equal-width classes of 10 minutes, make a frequency distribution, and describe the pattern that the histogram would show.
- Identify the observationsThe observational units are the eight students, and the quantitative variable is one-way commute time, measured in minutes. The sample size is eight. A frequency histogram is suitable for summarizing these numerical measurements.
- Set the classesThe smallest time is 12 minutes and the largest is 44 minutes, so classes from 10 up to 50 cover every value. Use equal-width classes of 10 minutes and include each lower endpoint but not the upper endpoint. Thus, a time of 20 belongs in the 20-to-under-30 class.
- Tally the observationsPlace 12 and 18 in the first class; 21, 23, 24, and 27 in the second; 31 in the third; and 44 in the last. Counting each value once gives frequencies of 2, 4, 1, and 1. The total is eight, matching the number of students.
- Calculate relative frequenciesDivide each class frequency by the sample size. For example, the second class contains four of the eight commute times. The percentages are included to show the share in each class; retain the exact simple fractions during calculation and report percentages to one decimal place.
- Interpret the displayDraw four touching bars over the stated intervals, with heights 2, 4, 1, and 1. The tallest bar is from 20 to under 30 minutes, so half of these students reported times in that interval. Most values are between 10 and under 30 minutes, while 44 minutes lies farther toward the high end. This small sample suggests a longer high-time tail, but the graph alone does not explain why the commute times differ.
Answer: The frequency distribution is 2, 4, 1, and 1 across the classes 10–under 20, 20–under 30, 30–under 40, and 40–50 minutes. The corresponding relative frequencies are 25.0%, 50.0%, 12.5%, and 12.5%.
Check: The class frequencies total 8, and the relative frequencies total 1.00, or 100.0%. Every observation is included exactly once.
Common mistakes and how to avoid them
Using overlapping classes so a boundary observation could be counted twice.
Correction: State a clear endpoint rule, such as including the lower endpoint and excluding the upper endpoint, and apply it consistently.
Leaving spaces between bars in a histogram.
Correction: Histogram bars for neighboring numerical classes touch. Gaps can suggest numerical intervals containing no observations.
Forgetting the axis labels or units.
Correction: Label the horizontal axis with the measured variable and its units, and the vertical axis with frequency or relative frequency.
Treating an unusual value as definitely incorrect.
Correction: Check the original record. An unusual value may be genuine, and the graph alone cannot decide.
Lesson summary
- Identify the observational units, quantitative variable, and measurement units before organizing values.
- Choose a dot plot, stem-and-leaf display, or histogram based on the size of the data set and whether individual values need to remain visible.
- For a frequency distribution, define classes carefully, tally each observation once, and verify the total frequency.
- Interpret a graph by describing shape, centre, spread, and possible unusual values in the context of the variable.
Check your understanding
Question 1
A class contains 6 observations out of a total of 24. What is its relative frequency?
- 0.25, or 25%
- 0.40, or 40%
- 4, or 400%
- 18, or 1800%
Show answer and explanation
0.25, or 25%
Divide the class count by the total: . Multiplying by 100 gives 25%.
Question 2
Which display is generally most useful for showing the individual values in a small numerical data set?
- A dot plot
- A pie chart
- A map
- A list of category labels
Show answer and explanation
A dot plot
A dot plot places a mark for each observation on a numerical scale, keeping individual values visible.
Question 3
A frequency distribution has class counts 3, 5, and 4. How many observations were recorded?
- 12
- 9
- 20
- 3
Show answer and explanation
12
Add the class frequencies: . Each observation should be counted in exactly one class.
Key terms
- Observational unit
- The person or object on which a measurement is made.
- Variable
- A characteristic recorded for each observational unit.
- Quantitative data
- Numerical measurements or counts for which arithmetic is meaningful.
- Frequency
- The number of observations in a value or class.
- Relative frequency
- The share of all observations in a class, found by dividing its frequency by the total number of observations.
- Histogram
- A graph that uses adjoining bars to show frequencies for numerical intervals.
- Distribution
- The pattern of values in a data set, including how often values occur and how they are arranged.
Continue through MATH 215
View the complete Athabasca University MATH 215: Introduction to Statistics learning path
- 1.1 · Use basic statistical terms and notation
- 1.2 · Classify variables and types of data
- 1.3 · Distinguish populations, samples, experiments, and summation notation
- 1.4 · Organize and graph qualitative data
- 1.6 · Calculate and interpret measures of centre for ungrouped data
- 1.7 · Calculate and interpret dispersion for ungrouped data
About this lesson
Published by DoAssignment. This AI-assisted lesson follows Athabasca University MATH 215: Introduction to Statistics, study topic 1.5. It is a study resource, not an official curriculum publication.