Data Science Program

Module 3: Data Visualization

Learn how to explore, interpret, and communicate data through histograms, boxplots, scatterplots, pie charts, and principled chart selection.

📖

Intermediate Level

150–180 Minutes

📊

Data Visualization

📅

Updated July 2026

Introduction

Statistics is often described as the science of learning from data. Before statistical models are fitted, hypotheses are tested, or Machine Learning algorithms are trained, analysts typically perform one essential task: they visualize the data. A well-designed graph can reveal patterns, trends, clusters, relationships, unusual observations, and potential errors that may remain hidden within numerical tables. Consequently, data visualization has become an indispensable component of modern statistics, business analytics, artificial intelligence, and data science.

Human beings naturally recognize visual patterns much more effectively than long lists of numbers. Consider a dataset containing one hundred customer purchase amounts. Although it is possible to inspect every numerical value individually, it is much easier to understand the overall distribution by viewing a histogram or a boxplot. Similarly, relationships between two numerical variables often become immediately apparent when displayed on a scatterplot, whereas they may be difficult to detect from a spreadsheet alone. Effective visualizations therefore complement numerical summaries by providing intuitive insights into the underlying structure of the data.

In professional practice, data visualization serves two important purposes. The first is exploration, where analysts examine the characteristics of a dataset before selecting appropriate statistical techniques. This process is commonly known as Exploratory Data Analysis (EDA). During exploratory analysis, visualizations help identify skewed distributions, unusual observations, missing values, potential data entry errors, and relationships between variables. The second purpose is communication. After an analysis has been completed, charts and graphs provide an effective way of presenting findings to managers, researchers, policymakers, and other stakeholders who may not possess advanced statistical knowledge.

Modern artificial intelligence systems also depend heavily upon data visualization during model development. Data scientists routinely visualize datasets before training predictive models to ensure that variables are distributed appropriately, relationships are reasonable, and anomalies are understood rather than ignored. A machine learning model trained on poorly understood data may produce misleading or unreliable predictions regardless of the sophistication of the algorithm. Consequently, visualization represents one of the most important quality assurance steps in the entire analytical workflow.

Throughout this module, you will learn how different graphical techniques summarize different aspects of a dataset. Some visualizations describe the distribution of a single variable, while others compare groups or illustrate relationships between two numerical variables. Rather than memorizing graph types, you will learn to select the visualization that best answers a particular analytical question. This analytical mindset closely reflects the workflow used by professional statisticians, business analysts, and data scientists.

💡Module Note: Data visualization does not replace statistical analysis. Instead, it provides the first opportunity to understand the structure and quality of the data before numerical summaries and statistical models are applied. Throughout this course, visualization and numerical analysis will complement one another, producing a more complete understanding of every dataset.

Learning Objectives

After completing this module, you should be able to:

  • Explain the importance of data visualization in statistics, artificial intelligence, and data science.
  • Select appropriate graphical techniques for different types of variables and analytical questions.
  • Interpret histograms, boxplots, scatterplots, bar charts, and pie charts.
  • Identify patterns, distributions, outliers, and relationships through graphical analysis.
  • Generate and interpret visualizations using the LearnerBox Interactive Dataset.
  • Recognize how visual exploration prepares data for correlation, regression, statistical inference, and machine learning.

Customer Purchase Analytics Dataset

The following dataset contains 150 customer records and includes variables relating to customer characteristics and purchasing activity.

Use the tabs below to move between the complete dataset, the data dictionary, numerical summaries, categorical frequency tables, and interactive visualisations.

LearnerBox Interactive Dataset

Dataset A: Customer Purchase Analytics

150 records 19 variables
Customer IDAgeGenderRegionEducationEmployment StatusLocationMarital StatusCard LoyaltyIncomePurchase AmountCredit ScoreSavings BalanceSatisfaction PreSatisfaction PostResponseTime PreResponseTime PostEntertainmentSessions PreEntertainmentSessions Post
CP00190Male121101110320977739292165.986.5410.38.1414
CP00263Female41111084482685682339275.757.398.96.4413
CP00396Female31111093518671655313866.16.1110.68.1314
CP00484Female231100117772792649303146.027.6312.610.6413
CP00599Male131011106703868710241435.096.618.74.9012
CP00649Male34101092826768685144666.387.588.74.216
CP00782Male341111117682958666339956.226.8212.77.229
CP00853Female22111186744826643295736.336.4110.96.7111
CP00959Female42101086331704657195395.516.28.24.8017
CP01071Female32111199112878616345965.825.8410.46.5214
CP01190Male131110111682746726312656.347.198.26.7210
CP01299Female221101102734874768280146.637.088.24.2113
CP01382Female131101220000872650314914.95.159.75.7011
CP01454Male331010102060709603151736.256.9251.5314
CP01554Male13000076049617536186465.395.475.42.726
CP01679Male3301118115775958095267.818.8283.628
CP01740Male32111093919707540282968.879.547308
CP01829Female45110099743715508305295.396.549.98.4010
CP01949Male24111096302686591255895.75.387.21.7416
CP02043Male41110074584474504245444.653.926.11.5016
CP02144Female141111102069927601338095.445.912.48.417
CP02239Male43100093758687503144066.047.968.76.929
CP02324Male14111093973720548262576.115.619.33.5410
CP02422Male24110096261537569312376.017.216.14.918
CP02599Male331011107399992633212976.357.983.217
CP02655Female11100089197659602180776.048.059.42.6014
CP02751Male12111188351848627318245.336.246.74013
CP02875Female331110110617722622287925.566.93106.805
CP02938Female33110193825870629251186.946.977.93.7310
CP03096Male321110106883682723351935.326.598.66.229
CP03183Female31110094867664627285194.635.577.95.9216
CP03228Male24000073743665500116225.826.613.39.2111
CP03354Male251111116906925663356765.315.689.27.6211
CP03491Female231101117262968718333956.866.666.92.329
CP03596Female11111097436751708311126.247.19118.209
CP03658Female42110194633868567359595.877.36.51.519
CP03770Male31101186727792618137436.697.559.77.7311
CP03834Male251010113878788538211727.287.964.92.3219
CP03998Female21110094371813683296098.24106.32.2413
CP04039Male431100101492705544309557.288.659.67.2012
CP04196Male1211101041967326703500565.8910.86.5013
CP04272Male33110197037888687312281.041.39.63.9013
CP04369Male231111103658924688319817.658.468.9808
CP04461Female341111104951955656337135.225.28.54.3111
CP04539Male12111088662676511277726.325.755.92.6212
CP04635Female41011046680450044931516.066.797.62.628
CP04738Male41111173943802572242606.876.368.25.858
CP04835Female35111197773802571238796.195.3411.510.217
CP04983Female341101113567953785375265.215.17105.339
CP05048Male43110183685731579250592.773.525.31.9414
CP05125Female33110179701766516273075.696.599.95.6210
CP05298Female131010116937822717178646.536.958.44.9310
CP05391Female31111098676730684307906.686.717.75.6211
CP05495Female3511011301051042743354586.17.2910.57.218
CP05595Female231101109023996746334847.18.0911.57.1413
CP05634Female32111082631656571265145.65.5911.29.2111
CP05752Male43110097077674566292077.387.997.41.5312
CP05881Male41110185439826579293924.424.58.54.916
CP05976Male21100097006745572159556.496.69.98.517
CP06033Male141100102696653590350925.936.3911.37.5315
CP06186Male41110091739689610306476.086.578.16.4012
CP06238Male4301116741972342599475.876.838.84.8011
CP06338Female12110185309782562305206.196.376.82.3221
CP06463Male42110082827683691318046.026.486.63.608
CP06568Male42101193367810621168806.647.8116.7114
CP06667Female22111090833681648251937.728.3112.27.6217
CP06728Male34101091554650541173554.854.9394016
CP06891Female221001101164877679151636.286.8911.45.9210
CP06942Male131101102426808624286574.244.31107.3512
CP07031Female23111192831763535310436.216.757.84.8415
CP07192Female431111108030990709345184.684.28107.2115
CP07297Male150110106622691719117264.026.0512.69.6112
CP07389Male241100125137777701334417.617.179.16.2010
CP07485Female221100105290639720356367.248.9212.68.2012
CP07573Female21110088645632615295874.874.9741.5417
CP07665Female42011065190480528153241.642.48116.9212
CP07727Female23101191927912559178905.995.879.45.4114
CP07853Male21111083405668536230594.934.128.53.7116
CP07950Male33001175912696553148527.697.99.46.7312
CP08041Male2301017144982052492897.227.964.81.5418
CP08178Female121101103504937706342205.145.098.73.9314
CP08260Female241100105209730649325415.365.08107.1011
CP08327Male11110078857470510274718.288.7812.59.8315
CP08438Male11101077154701562110436.567.68.14.4212
CP08575Male24000098508624702142247.27.457.43.7118
CP08627Male23100184686692544129186.487.278.14.2114
CP08789Female431100118514864714301107.318.2511.27.6416
CP08860Female231110100854847150318526.045.689.97210
CP08959Female32101188200881551208306.916.3786.158
CP09068Male341010108168770639190017.046.510.38.3215
CP09186Male141001115490958679149627.257.639.47015
CP09248Female33110197860902584311205.685.0911.36.2313
CP09375Female12111199599906650335509.29.197.74.4411
CP09466Female43100098972729650150685.776.588.57.1220
CP09542Female12110085356652551281346.717.48.34.107
CP09691Male2110111053138747111757611.6661.5617
CP09784Female151010132966848736179425.566.639.98114
CP09873Male41110190619790548310072.992.59.86.629
CP09988Male251011126007896691192385.196.029.26.1210
CP10090Male331001110885900695193489.14108.76.2011
CP10177Male11111092902616591266286.76.598.195.692.29.2
CP10287Female331000104323687707161205.815.848.866.862.989.98
CP10376Male141111112243921737333926.916.996.172.17-0.044.96
CP10475Female441010104702814707163528.127.5511.048.44-0.1615.84
CP10570Female141100113318833657365135.686.8410.779.572.598.59
CP10681Male13001085313549691123335.685.149.026.322.259.25
CP10733Male13111188026741547294778.27.87.924.722.1812.18
CP10891Female2411111265351050721354527.077.766.023.022.335.33
CP10990Male32100096283633648123745.354.997.423.520.7913.79
CP11023Male13110079498587492226686.767.3210.458.952.168.16
CP11187Male321101102246889661256485.365.448.564.162.2517.25
CP11231Male251011110844856589218735.3658.136.230.7413.74
CP11347Female251010119803825640173256.346.529.977.074.613.6
CP11480Female21111088524596636365273.354.211.279.872.5215.52
CP11533Female241111103895860591286973.613.569.496.490.038.03
CP11686Female35011199711885718115255.225.39.66.32.7911.79
CP11787Female11111192437753696338704.65.5511.698.390.3517.35
CP11861Male13011179628771521149826.447.399.336.832.9910.99
CP11921Female451100105441747450367294.745.118.373.773.549.54
CP12043Female41101070156599564176414.045.5610.318.910.589.58
CP12170Female23001185657810555166388.047.938.473.473.2510.25
CP12299Female2210011123741001713200155.697.068.252.752.4315.43
CP12371Male351111116563940627339926.16.711.59.53.0417.04
CP12441Male251101111254885532294724.023.588.43.44.6512.65
CP12571Male151100122223818708297685.257.283.72-1.781.4411.44
CP12664Male131100102423740662302336.166.619.665.060.686.68
CP12783Female341111113085938668307984.414.218.193.690.487.48
CP12829Female34100193638841441150056.536.98.415.610.5911.59
CP12981Male12111199182811638320085.174.329.876.371.6913.69
CP13064Female141111115020912672343685.65.18.657.052.329.32
CP13197Male11111199682779634302345.176.139.472.2211.22
CP13248Male33100093720757589208058.588.224.970.873.0511.05
CP13371Male12111099572636640265335.996.857.114.211.8311.83
CP13495Male11111081855575676276054.545.3711.769.563.9819.98
CP13529Female31110078615578518257117.157.138.947.441.4114.41
CP13695Male1300107854463860872554.315.15.660.565.8815.88
CP13775Female32111095449683689292976.36.686.850.952.758.75
CP13826Female14101090302734525156653.284.237.015.110.5311.53
CP13945Male11101071954527531187364.165.329.185.780.2110.21
CP14077Female121110103633753623322566.287.118.122.322.5314.53
CP14126Female33111085548606504254837.036.527.675.671.485.48
CP14267Male441000112034868622237666.245.438.997.592.8812.88
CP14348Male42111187451728563324095.845.859.924.222.529.52
CP14470Male231110105771745719323245.596.158.475.371.716.7
CP14524Male34100095665711484158883.955.477.842.340.5410.54
CP14690Male130010789896165731002154.468.143.74-0.4611.54
CP14798Female221001110240912775169085.376.049.657.651.1420.14
CP14898Female41101094660687676197337.478.388.965.163.0916.09
CP14980Male14011182208764598148756.487.1710.582.1310.13
CP15031Male11111071502551545317803.553.766.74.5-0.055.95

Why Data Visualization Matters

Imagine that you are presented with a spreadsheet containing one hundred customer records. Each record includes variables such as annual income, purchase amount, website visits, customer satisfaction, and product category. Although every value is available, it is extremely difficult to recognise meaningful patterns simply by scanning rows of numbers. A histogram immediately reveals the shape of a distribution, a boxplot highlights unusual observations, and a scatterplot can uncover relationships between variables that may otherwise remain unnoticed. Effective data visualization therefore transforms large collections of numbers into meaningful visual information that supports statistical reasoning and informed decision-making.

Visual exploration represents the first stage of Exploratory Data Analysis (EDA), a systematic approach to understanding the characteristics of a dataset before performing formal statistical analyses. Rather than beginning with formulas or hypothesis tests, analysts first ask simple but important questions. What does the distribution look like? Are there unusually large or unusually small observations? Do two variables appear to move together? Are there groups that differ substantially from one another? These questions help determine which statistical techniques are appropriate and whether the data satisfy the assumptions required for those techniques.

Exploratory Data Analysis was popularised by the statistician John Tukey, who argued that analysts should first "listen to the data" before applying mathematical models. Although modern statistical software performs sophisticated calculations almost instantly, the underlying principle remains unchanged. Careful visual inspection frequently identifies patterns, inconsistencies, or data quality issues that would otherwise influence the validity of subsequent analyses. Consequently, exploratory visualization has become a standard practice in statistics, business analytics, artificial intelligence, and machine learning.

Conceptual diagram illustrating the analytical workflow: Raw Data → Data Visualization → Statistical Insight → Statistical Analysis → Decision Making → Artificial Intelligence Applications. Highlight that visualization is the bridge between raw data and statistical reasoning.
Figure 3.1. Conceptual diagram illustrating the analytical workflow: Raw Data → Data Visualization → Statistical Insight → Statistical Analysis → Decision Making → Artificial Intelligence Applications. Highlight that visualization is the bridge between raw data and statistical reasoning.

As illustrated in Figure 3.1, visualization acts as an important bridge between raw data and statistical analysis. Before calculating descriptive statistics, estimating regression models, or training machine learning algorithms, analysts first seek to understand the overall structure of the data. Visualizations often reveal characteristics that influence subsequent statistical decisions, including skewed distributions, clusters of observations, outliers, and potential measurement errors.

Worked Example 3.1

Problem

Suppose an analyst receives the LearnerBox Customer Purchases Dataset (LB-CP-100 V1). Before calculating descriptive statistics or fitting predictive models, what questions should be explored through data visualization?

Solution

A systematic exploratory analysis might include questions such as:

  • How are customer purchase amounts distributed?
  • Are there unusually large purchases that may represent outliers?
  • Do annual income and purchase amount appear to be related?
  • Which product categories account for the largest proportion of purchases?
  • Does customer satisfaction vary substantially across different regions?

Each of these questions can be answered more effectively using appropriate graphical techniques than by examining numerical tables alone.

Interpretation

This example illustrates an important principle that will be followed throughout the remainder of this course: visual interpretation should generally precede numerical interpretation. Graphs provide an initial understanding of the data, allowing analysts to identify interesting patterns before computing statistical summaries or fitting analytical models.

🤖AI Connection: Professional data scientists rarely begin a machine learning project by training a predictive model immediately. Instead, they first visualise the available data to identify missing values, unusual observations, skewed variables, and relationships between predictors. This exploratory process improves data quality and frequently leads to better-performing models.
🧭Try it Yourself: Using the LearnerBox Interactive Dataset, select the Histogram chart type and choose the variable Purchase Amount. Without calculating any numerical statistics, examine the graph and consider the following questions:
  • Does the distribution appear approximately symmetric or skewed?
  • Are there unusually large purchases?
  • Would the mean likely be greater than, less than, or approximately equal to the median?

Do not worry if you cannot answer every question immediately. By the end of this module, you will be able to interpret each of these characteristics confidently.

Notice that every visualization introduced throughout this module answers a different analytical question. Histograms describe distributions, boxplots summarise variability and identify potential outliers, scatterplots investigate relationships between numerical variables, while pie charts illustrate the relative proportions of categorical variables. Selecting the appropriate visualization therefore depends not on personal preference but on the specific analytical question being investigated.

This principle will guide the remainder of this module. Rather than learning graph types in isolation, you will learn to choose the visualization that best communicates the information contained within the data.

Histograms

Analytical Question: What does the distribution of a numerical variable look like?

Among all statistical graphs, the histogram is one of the most widely used tools for exploring numerical data. Unlike tables of numbers, which often conceal important characteristics, a histogram provides an immediate visual summary of how observations are distributed across a range of values. It enables analysts to identify the centre of the distribution, the overall spread of the data, the presence of multiple peaks, and whether the distribution appears approximately symmetric or skewed. Consequently, histograms are almost always one of the first graphs produced during exploratory data analysis.

A histogram is constructed by dividing the range of a numerical variable into consecutive intervals, known as bins, and counting the number of observations that fall within each interval. The horizontal axis represents the values of the variable, while the vertical axis represents the corresponding frequencies. Because the bins represent continuous intervals, the bars in a histogram are adjacent to one another, illustrating that the underlying variable is quantitative rather than categorical. This distinguishes histograms from bar charts, whose bars are separated because they represent discrete categories.

Histograms provide valuable information about the overall shape of a distribution. For example, a histogram may reveal that the observations are approximately normally distributed, concentrated around a central value with similar frequencies on either side. Alternatively, the graph may indicate that the distribution is positively skewed, with a long tail extending toward larger values, or negatively skewed, with the tail extending toward smaller values. Histograms can also reveal multiple peaks (multimodality), unusually large gaps between observations, or isolated values that may warrant further investigation.

One important characteristic of a histogram is that it summarises the distribution without displaying individual observations. Although this makes the graph easier to interpret, it also means that the appearance of the histogram depends partly on the number and width of the bins selected. Too few bins may oversimplify the distribution, while too many bins may exaggerate minor fluctuations caused by random variation. Modern statistical software automatically selects sensible bin widths, although analysts should always examine whether the resulting histogram provides an informative representation of the data.

Histogram of Income generated directly from the LearnerBox Customer Purchases Dataset (LB-CP-100 V1).
Figure 3.2. Histogram of Annual Income generated directly from the LearnerBox Customer Purchases Dataset (LB-CP-100 V1).

Figure 3.2 presents the distribution of Annual Income for customers contained within the LearnerBox Customer Purchases Dataset. Even before calculating descriptive statistics, several important observations can be made from the histogram. Most customers fall within the middle income ranges, while comparatively few occupy the lowest and highest income categories. The graph also allows us to determine whether the distribution appears approximately symmetric or exhibits noticeable skewness. Such observations provide useful context for interpreting subsequent numerical summaries such as the mean, median, standard deviation, and quartiles.

Worked Example 3.2

Problem

Use the LearnerBox Interactive Dataset to examine the distribution of Annual Income. Describe the overall shape of the histogram and identify any notable characteristics.

Solution

Select the following options within the LearnerBox Interactive Dataset Viewer.

  • Chart Type: Histogram
  • Variable: Annual Income

Observe the resulting histogram carefully before calculating any descriptive statistics. Consider the following questions:

  • Where are most observations concentrated?
  • Does the distribution appear approximately symmetric?
  • Are there any unusually high or low values?
  • Would a normal distribution provide a reasonable approximation?

Interpretation

The histogram indicates that annual incomes are concentrated within the middle ranges, with relatively fewer customers occupying the extreme income levels. Although the precise numerical characteristics will be examined in later modules, the graph provides an immediate overview of the distribution and helps determine whether subsequent statistical techniques based on normality are likely to be appropriate.

🧭Try it Yourself: Using the LearnerBox Interactive Dataset Viewer, generate histograms for the following variables:

  • Purchase Amount
  • Credit Score

Compare the three histograms. Which variable appears most symmetric? Which appears most skewed? Which variable shows the greatest variability? Try answering these questions using only the visualisations before examining any numerical summaries.

Histogram of Purchase Amount generated directly from the LearnerBox Customer Purchases Dataset (LB-CP-100 V1).
Figure 3.3. Histogram of Purchase Amount generated directly from the LearnerBox Customer Purchases Dataset (LB-CP-100 V1).

Figure 3.3 illustrates the distribution of Purchase Amount. Compared with Annual Income, the purchase amounts exhibit greater variability and may display a longer right-hand tail resulting from a small number of customers making unusually large purchases. Such positively skewed distributions are frequently encountered in retail and e-commerce data because relatively few customers account for a disproportionately large share of total sales. Visual inspection therefore helps analysts anticipate the presence of high-value observations before formal statistical analyses are undertaken.

Histogram of Credit Score generated directly from the LearnerBox Customer Purchases Dataset (LB-CP-100 V1).
Figure 3.4. Histogram of Credit Score generated directly from the LearnerBox Customer Purchases Dataset (LB-CP-100 V1).

Figure 3.4 presents the distribution of Credit Score. Unlike Purchase Amount, this variable reflects customer card usage behaviour rather than financial activity. Examining the histogram allows analysts to determine whether most customers visiting a website credit score differences exist between casual visitors and highly engaged users. Such information becomes valuable in later modules when investigating relationships between browsing behaviour and purchasing decisions using correlation and regression analysis.

Histograms demonstrate one of the fundamental principles of exploratory data analysis: a single graph can often communicate more information than an entire page of numerical values. Throughout the remainder of this module, you will discover that different graphical techniques reveal different aspects of the same dataset. The next visualization, the boxplot, complements the histogram by providing a compact summary of the distribution while simultaneously highlighting variability and potential outliers.

Boxplots

Analytical Question: How are the data distributed, and are there any unusual observations or differences between groups?

While histograms provide a detailed picture of the overall shape of a distribution, analysts often require a more compact summary that highlights the centre, spread, and potential outliers of the data. The boxplot, sometimes called a box-and-whisker plot, was developed specifically for this purpose. Despite its simple appearance, the boxplot conveys a considerable amount of statistical information and has become one of the most widely used visualisation tools in exploratory data analysis.

A boxplot is constructed from the five-number summary introduced in Module 2. These five statistics consist of the minimum value, the first quartile (Q1), the median (Q2), the third quartile (Q3), and the maximum value. The box itself represents the interquartile range (IQR), extending from the first quartile to the third quartile and therefore containing the middle fifty percent of the observations. A horizontal line inside the box marks the median, providing an immediate indication of the centre of the distribution.

The lines extending from the box are known as whiskers. In most statistical software, the whiskers extend to the most extreme observations that are not considered unusual according to the 1.5 × IQR rule. Observations beyond these limits are displayed individually as points and are interpreted as potential outliers. It is important to recognise that an outlier is not necessarily an error. Instead, it is an observation that differs substantially from the majority of the data and therefore deserves further investigation before any decision is made about retaining or removing it.

One of the greatest strengths of the boxplot is its ability to compare several groups simultaneously. Unlike histograms, which become difficult to interpret when multiple distributions are displayed together, grouped boxplots allow analysts to compare medians, variability, skewness, and potential outliers across different categories in a clear and compact manner. For this reason, grouped boxplots are frequently used before conducting statistical procedures such as the independent-samples t-test or analysis of variance (ANOVA).

Boxplot of Purchase Amount generated from the LearnerBox Customer Purchases Dataset (LB-CP-100 V1).
Figure 3.5. Boxplot of Purchase Amount generated from the LearnerBox Customer Purchases Dataset (LB-CP-100 V1).

Figure 3.5 presents a boxplot of Purchase Amount from the LearnerBox Customer Purchases Dataset. Notice how the graph immediately summarises the centre of the distribution through the median, illustrates the spread of the middle fifty percent of observations using the interquartile range, and identifies unusually large purchases that appear as individual points beyond the whiskers. Information that would require several numerical statistics can therefore be communicated in a single compact figure. The boxplot also shows several outliers, which means data points that are more than 1.5 time the sample IQR value.

Worked Example 3.3

Problem

Use the LearnerBox Interactive Dataset to investigate the distribution of Purchase Amount using a boxplot.

Solution

Within the LearnerBox Interactive Dataset Viewer, select:

  • Chart Type: Boxplot
  • Variable: Purchase Amount

Examine the resulting graph and answer the following questions.

  • Where is the median located?
  • How wide is the interquartile range?
  • Are any potential outliers visible?
  • Does the distribution appear approximately symmetric or skewed?

Interpretation

The boxplot provides an immediate summary of the distribution while drawing attention to observations that deserve further investigation. Compared with a histogram, the boxplot sacrifices some detail about the exact shape of the distribution but provides a much clearer summary of variability and potential outliers.

Grouped boxplots comparing Purchase Amount across the five geographical regions represented in the LearnerBox Customer Purchases Dataset.
Figure 3.6. Grouped boxplots comparing Purchase Amount across the five geographical regions represented in the LearnerBox Customer Purchases Dataset.

Grouped boxplots extend the usefulness of this visualisation by allowing several distributions to be compared simultaneously. In Figure 3.6, Purchase Amount is displayed separately for each geographical region. Rather than comparing dozens of individual observations, the analyst can immediately compare the medians, variability, and potential outliers for all four groups. Such visual comparisons frequently provide valuable insights before any formal statistical hypothesis tests are performed.

Conceptual comparison of several representative boxplots illustrating differences in median, variability, skewness, and outliers.
Figure 3.7. Conceptual comparison of several representative boxplots illustrating differences in median, variability, skewness, and outliers.

Figure 3.7 should be viewed as a general interpretation guide rather than the summary of a particular dataset. The purpose of this figure is to demonstrate how different characteristics of a distribution influence the appearance of a boxplot. By comparing several representative boxplots side by side, students can learn to recognise important distributional features without becoming distracted by the underlying numerical values.

The following characteristics are particularly important when interpreting boxplots.

  • Median: A higher median indicates that the typical observations occur at larger values.
  • Interquartile Range (IQR): A wider box indicates greater variability among the middle fifty percent of the observations.
  • Whisker Length: Unequal whiskers often suggest skewness in the distribution.
  • Outliers: Individual points beyond the whiskers identify observations that differ markedly from the remainder of the data.
  • Comparisons Between Groups: Differences in medians, variability, and outliers often suggest meaningful differences that may later be examined using inferential statistical techniques.
❗Important Note: A boxplot should never be interpreted in isolation. Whenever possible, examine both the histogram and the corresponding boxplot. The histogram provides a detailed view of the distribution's shape, whereas the boxplot summarises the same distribution using the five-number summary and highlights potential outliers. Together, these visualisations provide a much more complete understanding of the data than either graph alone.
🧭Try it Yourself: Using the LearnerBox Interactive Dataset, compare grouped boxplots of Purchase Amount by Region and by Education.

As you compare the graphs, consider the following questions:

  • Which group has the highest median purchase amount?
  • Which group displays the greatest variability?
  • Are potential outliers present in one group or several groups?
  • Would you expect the group means to differ substantially? Why or why not?

You will revisit these same grouped boxplots in Module 6 when studying Analysis of Variance (ANOVA), where visual observations will be supported by formal statistical tests.

Scatterplots

Analytical Question: Are two numerical variables related to one another?

Many statistical investigations seek to determine whether changes in one numerical variable are associated with changes in another. For example, do customers with higher annual incomes tend to spend more money? Does the amount of time spent on a website influence purchase behaviour? Do students who study for more hours generally achieve higher examination scores? Questions such as these cannot be answered effectively using histograms or boxplots because those visualisations describe only one variable at a time. Instead, analysts use the scatterplot, one of the most informative graphical techniques in statistics.

A scatterplot displays the relationship between two numerical variables by representing each observation as a single point on a two-dimensional graph. The horizontal axis represents the explanatory variable (also called the independent or predictor variable), while the vertical axis represents the explained variable (also called the dependent or response variable). Rather than grouping observations into intervals, every data point is plotted individually, allowing the overall pattern of the relationship to emerge naturally.

Scatterplots provide considerably more information than simply indicating whether two variables are related. They help analysts assess the direction of the relationship, its strength, its form, and the possible presence of outliers. These four characteristics form the foundation of visual relationship analysis and will be explored quantitatively in the following modules.

The direction of the relationship indicates whether the variables tend to increase together or move in opposite directions. If larger values of one variable are generally associated with larger values of the other variable, the relationship is described as positive. Conversely, if larger values of one variable are generally associated with smaller values of the other variable, the relationship is described as negative. When no consistent pattern is apparent, the variables may exhibit little or no relationship.

The strength of the relationship refers to how closely the observations cluster around an underlying trend. When the points lie close to an imaginary straight line, the relationship is considered strong. When the points are widely scattered, the relationship is weaker. Importantly, visual inspection provides only an initial impression of relationship strength. In Module 4, this visual assessment will be quantified using the Pearson correlation coefficient, while Module 6 will extend the analysis through regression modelling.

Another important characteristic is the form of the relationship. Many statistical techniques assume that the association between two variables is approximately linear. A scatterplot allows analysts to evaluate this assumption before fitting statistical models. If the relationship follows a curved pattern, alternative analytical techniques may be more appropriate.

Finally, scatterplots frequently reveal outliers that differ markedly from the general pattern. These observations may represent data entry errors, unusual cases, or genuinely informative observations. Because influential observations can substantially affect regression models, they should always be investigated carefully rather than removed automatically.

Scatterplot illustrating the relationship between Annual Income and Purchase Amount using the LearnerBox Customer Purchases Dataset (LB-CP-100 V1).
Figure 3.8. Scatterplot illustrating the relationship between Annual Income and Purchase Amount using the LearnerBox Customer Purchases Dataset (LB-CP-100 V1).

Figure 3.8 illustrates the relationship between Annual Income and Purchase Amount. Each point represents one customer from the LearnerBox Customer Purchases Dataset. Although some variability is expected, the overall pattern suggests that customers with higher annual incomes generally tend to spend more money. At the same time, the graph demonstrates that annual income alone does not completely determine purchasing behaviour, highlighting the influence of additional explanatory variables such as previous purchases, loyalty membership, promotional discounts, and customer preferences.

Worked Example 3.4

Problem

Use the LearnerBox Interactive Dataset to investigate the relationship between Annual Income and Purchase Amount.

Solution

Within the LearnerBox Interactive Dataset Viewer, select:

  • Chart Type: Scatterplot
  • X Variable: Annual Income
  • Y Variable: Purchase Amount

Before calculating any numerical statistics, examine the scatterplot carefully.

  • Does the relationship appear positive or negative?
  • Is the relationship relatively strong or relatively weak?
  • Does the relationship appear approximately linear?
  • Can you identify any unusual observations?

Interpretation

The scatterplot suggests a positive relationship between annual income and purchase amount. Although considerable variation exists, customers with larger annual incomes generally tend to make larger purchases. This visual impression will later be confirmed quantitatively through correlation and regression analysis.

Scatterplot illustrating the relationship between Age and Credit Scores using the LearnerBox Customer Purchases Dataset (LB-CP-100 V1).
Figure 3.9. Scatterplot illustrating the relationship between Age and Credit Card Scores using the LearnerBox Customer Purchases Dataset (LB-CP-100 V1).

Unlike the previous example, Figure 3.9 compares a customer behavior variable and a demographic variable. The graph demonstrates how scatterplots can also be used to investigate behavioural relationships rather than purely economic ones. As customer age increases, the inclination to spend more using credit cards generally increases as well, illustrating another positive association that can later be modelled statistically.

🤖AI Connection: Scatterplots are among the first visualisations examined by data scientists before developing predictive models. By exploring relationships between explanatory variables and target variables, analysts can identify potentially useful predictors, detect influential observations, and evaluate whether simple linear models are likely to provide reasonable approximations before applying more sophisticated machine learning algorithms.
🧭Try it Yourself: Using the LearnerBox Interactive Dataset Viewer, investigate each of the following relationships.

  • Bank Balance versus Age
  • Income versus Credit Score
  • Purchase Amount versus Credit Score

For each scatterplot, describe:

  • The direction of the relationship.
  • The apparent strength of the relationship.
  • Whether the relationship appears approximately linear.
  • Whether any unusual observations deserve further investigation.

Compare your visual conclusions with those of your classmates. Although different analysts may notice different features, the overall interpretation should be broadly consistent.

Scatterplots represent an important transition within this course. Until now, the emphasis has been on describing individual variables. Beginning with scatterplots, the focus shifts toward understanding relationships between variables, a theme that continues throughout the remainder of the textbook in correlation analysis, regression modelling, statistical learning, and artificial intelligence.

Pie Charts

Analytical Question: How is a categorical variable divided into its constituent parts?

Not all variables consist of numerical measurements. Many datasets contain categorical variables that classify observations into distinct groups, such as geographical regions, product categories, customer segments, educational qualifications, or preferred software applications. When the objective is to illustrate the relative proportion contributed by each category to the whole dataset, one of the most familiar graphical techniques is the pie chart. Although simple in appearance, pie charts remain widely used in business reports, dashboards, market research, and management presentations because they communicate proportions quickly to non-technical audiences.

A pie chart represents an entire dataset as a circle divided into sectors. Each sector corresponds to one category, and the size of each sector is proportional to either its frequency or its percentage of the total observations. Categories that occur more frequently occupy larger sectors, whereas less common categories occupy correspondingly smaller sectors. Unlike histograms and scatterplots, which display numerical relationships, pie charts are intended exclusively for categorical data.

Because the human eye finds it easier to compare lengths than angles, statisticians generally recommend using pie charts only when the number of categories is relatively small and when the purpose is to communicate broad proportions rather than precise numerical differences. When many categories are present or when accurate comparisons between categories are required, a bar chart usually provides a more effective visualisation. Consequently, analysts should select pie charts carefully rather than using them by default.

Pie chart illustrating the proportion of customers belonging to each Education category using the LearnerBox Customer Purchases Dataset (LB-CP-100 V1).
Figure 3.10. Pie chart illustrating the proportion of customers belonging to each Education category using the LearnerBox Customer Purchases Dataset (LB-CP-100 V1).

Figure 3.10 displays the distribution of customers across the various Education categories contained within the LearnerBox Customer Purchases Dataset. Rather than presenting the raw frequencies in a table, the pie chart immediately communicates the relative contribution of each category to the entire customer base. Managers can therefore identify dominant product segments at a glance without examining detailed numerical summaries.

Worked Example 3.5

Problem

Use the LearnerBox Interactive Dataset to determine the proportion of customers belonging to each education category.

Solution

Within the LearnerBox Interactive Dataset Viewer, select:

  • Chart Type: Pie Chart
  • Variable: Education

Observe the completed chart and answer the following questions.

  • Which category occupies the largest sector?
  • Which category contributes the smallest proportion?
  • Are the categories distributed relatively evenly, or does one category dominate?
  • Would a bar chart communicate the same information more effectively?

Interpretation

The pie chart provides an immediate overview of the composition of the customer population. Although exact percentages can be obtained from numerical frequency tables, the visual representation makes it considerably easier to appreciate the relative importance of each category within the dataset.

One important principle should always be remembered when interpreting pie charts. The chart illustrates relative proportions, not numerical relationships. It cannot indicate trends, variability, correlations, or distributions. Consequently, pie charts complement rather than replace other statistical visualisations. Their primary purpose is to answer a single question: How is the whole divided among its categories?

✅Good Practice: Pie charts work best when:

  • Only a small number of categories are present.
  • The categories represent parts of a whole.
  • The objective is to communicate proportions rather than precise comparisons.
  • Sector labels include percentages for easier interpretation.

If the dataset contains many categories or if accurate comparisons are required, consider using a bar chart instead.

🧭Try it Yourself: Using the LearnerBox Interactive Dataset Viewer, create pie charts for each of the following variables.

  • Region
  • Gender
  • Employment Status

After creating each chart, compare it with the corresponding bar chart.

Which visualisation allows you to compare the categories more accurately? Which visualisation communicates overall proportions more effectively? Reflect on the advantages and limitations of each graphical technique before proceeding to the next section.

Here are two more worked out examples for your practice.

Worked Example 3.6

Problem

A retail company wishes to understand whether the variable Credit Score is suitable for further statistical analysis. Before calculating descriptive statistics, the analyst decides to examine the distribution visually.

Using the LearnerBox Interactive Dataset, generate a histogram of Credit Score and interpret the distribution.

Solution

Within the LearnerBox Interactive Dataset Viewer, select:

  • Chart Type: Histogram
  • Variable: Website Visits

Carefully inspect the resulting histogram and consider the following questions.

  • Where are most observations concentrated?
  • Does the distribution appear approximately symmetric or positively skewed?
  • Are there unusually large numbers of website visits?
  • Would the mean likely be influenced by a small number of extreme observations?

Interpretation

The histogram demonstrates why visual exploration should precede numerical analysis. Even before calculating the mean or standard deviation, the analyst gains valuable information about the overall distribution of website visits. If the histogram exhibits substantial skewness or contains unusually large observations, these characteristics should be considered when selecting subsequent statistical techniques. Visual interpretation therefore provides an important context for understanding later numerical summaries and statistical models.

Worked Example 3.7

Problem

An online retailer believes that customers with higher credit scores tend to make larger purchases. Before calculating a correlation coefficient or fitting a regression model, the analyst first examines the relationship visually.

Using the LearnerBox Interactive Dataset, construct a scatterplot of Credit Score and Purchase Amount.

Solution

Within the LearnerBox Interactive Dataset Viewer, select:

  • Chart Type: Scatterplot
  • X Variable: Credit Score
  • Y Variable: Purchase Amount

Study the resulting scatterplot and answer the following questions.

  • Does the relationship appear positive, negative, or absent?
  • How closely do the observations cluster around the overall trend?
  • Does the relationship appear approximately linear?
  • Can you identify any observations that differ noticeably from the general pattern?

Interpretation

The scatterplot provides an initial visual assessment of the relationship between browsing behaviour and customer spending. If the points exhibit an upward linear trend, the analyst may reasonably anticipate a positive correlation between the variables. Conversely, a widely scattered pattern would suggest a weaker relationship. Importantly, the scatterplot also reveals unusual observations that may influence the strength of the relationship and should therefore be investigated before formal correlation or regression analyses are performed. This illustrates an important principle that will recur throughout the remainder of this course: every regression analysis should begin with a scatterplot.

At this point in the module you have encountered the principal graphical techniques used during exploratory data analysis. Histograms describe the distribution of numerical variables, boxplots summarise variability and identify potential outliers, scatterplots investigate relationships between numerical variables, while pie charts communicate the composition of categorical variables. The final section of this module brings these visualisations together by developing a practical framework for selecting the most appropriate chart for any analytical question.

Choosing the Correct Chart

One of the most common mistakes made by beginning analysts is to select a graph simply because it is familiar rather than because it is appropriate. Effective data visualization begins not with the graph itself but with a clear understanding of the analytical question. Once the objective of the analysis has been established, the choice of visualization becomes much more straightforward. Throughout this module, you have seen that each graphical technique highlights different characteristics of the data, and no single graph is suitable for every situation.

A useful way to approach visualization is to begin by identifying the type of variables involved. If the objective is to describe the distribution of a single numerical variable, a histogram is usually the most informative choice. If the goal is to summarise the distribution while identifying variability and potential outliers, a boxplot provides a more compact representation. When the objective is to investigate the relationship between two numerical variables, a scatterplot allows the analyst to examine direction, strength, linearity, and unusual observations. Finally, when the purpose is to communicate the relative proportions of a categorical variable, a pie chart provides an effective visual summary, particularly when only a small number of categories are present.

Selecting an appropriate graph is not merely a matter of presentation. The visualization chosen often determines how quickly patterns are recognised, whether assumptions are questioned, and whether important observations receive further investigation. Consequently, experienced analysts rarely ask, "Which graph should I draw?" Instead, they ask, "What question am I trying to answer?" The graph then becomes a tool for answering that question as clearly and accurately as possible.

Decision guide comparing common statistical visualizations and the analytical questions they answer.
Figure 3.11. Decision guide comparing common statistical visualizations and the analytical questions they answer.

Figure 3.11 summarises the principal visualization techniques introduced in this module and the situations in which they are most appropriately used.

Table 3.1 Choosing the Appropriate Statistical Visualization
Analytical Question Recommended Visualization
What does the distribution of a numerical variable look like? Histogram
Are there unusual observations or differences between groups? Boxplot
Are two numerical variables related? Scatterplot
How is a categorical variable divided into its constituent parts? Pie Chart
Is a numerical variable normally distributed? Q-Q Plot
Which visualization should I choose? Select the graph that best answers the analytical question rather than the one that appears most familiar.
✅Good Analytical Practice: Always begin with the analytical question rather than the graph. The purpose of visualization is to improve understanding, not simply to produce attractive figures. Well-designed visualizations communicate statistical information accurately, efficiently, and honestly.

Practice Exercises

The following exercises are designed to reinforce the concepts introduced throughout this module by encouraging you to explore the LearnerBox Interactive Dataset. Unless otherwise stated, use the embedded dataset viewer to generate each visualization before attempting to answer the accompanying questions.

  1. Construct a histogram of Annual Income.
    • Describe the overall shape of the distribution.
    • Does the distribution appear approximately symmetric or skewed?
    • Are there any unusually large or unusually small observations?
  2. Generate a boxplot of Purchase Amount.
    • Estimate the location of the median.
    • Describe the variability of the middle fifty percent of observations.
    • Identify any potential outliers.
  3. Produce grouped boxplots of Purchase Amount by Region.
    • Which region appears to have the highest median purchase amount?
    • Which region displays the greatest variability?
    • Would you expect a formal ANOVA to detect differences between the regions?
  4. Construct a scatterplot of Annual Income and Purchase Amount.
    • Describe the direction of the relationship.
    • Would you describe the relationship as weak, moderate, or strong?
    • Does the relationship appear approximately linear?
    • Can you identify any observations that may deserve further investigation?
  5. Create a pie chart of Education Status.
    • Which category occupies the largest proportion of the customer population?
    • Would a bar chart communicate the same information more effectively? Explain your answer.
  6. Using the same dataset, determine which visualization would be most appropriate for each of the following:
    • Investigating customer age distribution.
    • Comparing purchase amounts across regions.
    • Examining the relationship between age and purchase amount.
    • Determining the proportion of loyalty programme members.
📥Reflection Activity: Suppose you have been asked to prepare a short presentation for the management team of an online retail company. You have access to the LearnerBox Customer Purchases Dataset and may include only three visualizations in your presentation.

Which three graphs would you choose, and why?

There is no single correct answer. Your selection should depend on the questions you believe management would most likely wish to answer.

Module Summary

Data visualization represents one of the most important stages of statistical analysis because it transforms raw numerical information into meaningful visual patterns that can be interpreted quickly and effectively. Throughout this module, you learned that visual exploration should generally precede numerical analysis, allowing analysts to understand the characteristics of a dataset before calculating descriptive statistics or fitting statistical models. You examined the principal graphical techniques used in exploratory data analysis, including histograms, boxplots, scatterplots, and pie charts, and learned that each visualization answers a different analytical question. By working directly with the LearnerBox Interactive Dataset, you also experienced how modern statistical software allows analysts to investigate data dynamically rather than relying solely on static textbook examples. These graphical techniques provide the visual foundation upon which the remaining modules of this course will build, particularly in correlation analysis, regression modelling, hypothesis testing, and artificial intelligence.

Key Takeaways

  • Data visualization is the first step in exploratory data analysis and helps reveal patterns that may not be apparent from numerical summaries alone.
  • Histograms describe the distribution of a single numerical variable and help identify symmetry, skewness, and variability.
  • Boxplots summarise distributions using the five-number summary while highlighting potential outliers and facilitating group comparisons.
  • Scatterplots investigate relationships between two numerical variables and provide the visual foundation for correlation and regression analysis.
  • Pie charts communicate the relative proportions of categorical variables and are most effective when only a small number of categories are present.
  • The choice of visualization should always be guided by the analytical question rather than personal preference or familiarity.
  • Visual interpretation should precede numerical interpretation, just as conceptual understanding should precede software implementation.

You have successfully completed the module content. Ready to Test What You Learned?

Take a short Quiz and find your score. You can always come back to this page and go through the content again!