- Descriptive Statistics




 Statistics is concerned with the describing, interpretation and analyzing of data.  It is, therefore, an essential element in any improvement process.  Statistics is often categorized into descriptive and inferential statistics.  It uses analytical methods which provide the math to model and predict variation.  It uses graphical methods to help making numbers visible for communication purposes.

Why do we Need Statistics?  To find why a process behaves the way it does.  To find why it produces defective goods or services.  To center our processes on ‘Target’ or ‘Nominal’.  To check the accuracy and precision of the process.  To prevent problems caused by assignable causes of variation.  To reduce variability and improve process capability.  To know the truth about the real world.

Descriptive Statistics:  Methods of describing the characteristics of a data set.  Useful because they allow you to make sense of the data.  Helps exploring and making conclusions about the data in order to make rational decisions.  Includes calculating things such as the average of the data, its spread and the shape it produces.


For example, we may be concerned about describing: • The weight of a product in a production line. • The time taken to process an application.

Descriptive statistics involves describing, summarizing and organizing the data so it can be easily understood.  Graphical displays are often used along with the quantitative measures to enable clarity of communication.


When analyzing a graphical display, you can draw conclusions based on several characteristics of the graph.  You may ask questions such ask: • Where is the approximate middle, or center, of the graph? • How spread out are the data values on the graph? • What is the overall shape of the graph? • Does it have any interesting patterns?


Outlier:  A data point that is significantly greater or smaller than other data points in a data set.  It is useful when analyzing data to identify outliers  They may affect the calculation of descriptive statistics.  Outliers can occur in any given data set and in any distribution.


 The easiest way to detect them is by graphing the data or using graphical methods such as: • Histograms. • Boxplots. • Normal probability plots.

Outliers may indicate an experimental error or incorrect recording of data.  They may also occur by chance. • It may be normal to have high or low data points.  You need to decide whether to exclude them before carrying out your analysis. • An outlier should be excluded if it is due to measurement or human error.


The following measures are used to describe a data set:  Measures of position (also referred to as central tendency or location measures).  Measures of spread (also referred to as variability or dispersion measures).  Measures of shape.

 If assignable causes of variation are affecting the process, we will see changes in: • Position. • Spread. • Shape. • Any combination of the three.


Measures of Position:  Position Statistics measure the data central tendency.  Central tendency refers to where the data is centered.  You may have calculated an average of some kind.  Despite the common use of average, there are different statistics by which we can describe the average of a data set: • Mean. • Median. • Mode.


Measures of Position:  Position Statistics measure the data central tendency.  Central tendency refers to where the data is centered.  You may have calculated an average of some kind.  Despite the common use of average, there are different statistics by which we can describe the average of a data set: • Mean. • Median. • Mode.

Median:  The middle value where exactly half of the data values are above it and half are below it.  Less widely used.  A useful statistic due to its robustness.  It can reduce the effect of outliers.  Often used when the data is nonsymmetrical.  Ensure that the values are ordered before calculation.  With an even number of values, the median is the mean of the two middle values.

Mode:  The value that occurs the most often in a data set.  It is rarely used as a central tendency measure  It is more useful to distinguish between unimodal and multimodal distributions • When data has more than one peak


Measures of Spread:  The Spread refers to how the data deviates from the position measure.  It gives an indication of the amount of variation in the process. • An important indicator of quality. • Used to control process variability and improve quality.  All manufacturing and transactional processes are variable to some degree.  There are different statistics by which we can describe the spread of a data set: • Range. • Standard deviation.

Range:  The difference between the highest and the lowest values.  The simplest measure of variability.  Often denoted by ‘R’.  It is good enough in many practical cases.  It does not make full use of the available data.  It can be misleading when the data is skewed or in the presence of outliers. • Just one outlier will increase the range dramatically.

Standard Deviation:  The average distance of the data points from their own mean.  A low standard deviation indicates that the data points are clustered around the mean.  A large standard deviation indicates that they are widely scattered around the mean.  The standard deviation of a sample is denoted by ‘s’.  The standard deviation of a population is denoted by “μ”.\



andard Deviation:  Perceived as difficult to understand because it is not easy to picture what it is.  It is however a more robust measure of variability.  Standard deviation is computed as follows:


 Exercise:  This example is about the time taken to process a sample of applications.  Find the mean, median, range and standard deviation for the following set of data: 2.8, 8.7, 0.7, 4.9, 3.4, 2.1 & 4.0.


Measures of Shape:  Data can be plotted into a histogram to have a general idea of its shape, or distribution.  The shape can reveal a lot of information about the data.  Data will always follow some know distribution.


It may be symmetrical or nonsymmetrical.  In a symmetrical distribution, the two sides of the distribution are a mirror image of each other.  Examples of symmetrical distributions include: • Uniform. • Normal. • Camel-back. • Bow-tie shaped


The shape helps identifying which descriptive statistic is more appropriate to use in a given situation.  If the data is symmetrical, then we may use the mean or median to measure the central tendency as they are almost equal.  If the data is skewed, then the median will be a more appropriate to measure the central tendency.  Two common statistics that measure the shape of the data: • Skewness. • Kurtosis.


The shape helps identifying which descriptive statistic is more appropriate to use in a given situation.  If the data is symmetrical, then we may use the mean or median to measure the central tendency as they are almost equal.  If the data is skewed, then the median will be a more appropriate to measure the central tendency.  Two common statistics that measure the shape of the data: • Skewness. • Kurtosis.


The shape helps identifying which descriptive statistic is more appropriate to use in a given situation.  If the data is symmetrical, then we may use the mean or median to measure the central tendency as they are almost equal.  If the data is skewed, then the median will be a more appropriate to measure the central tendency.  Two common statistics that measure the shape of the data: • Skewness. • Kurtosis.


 Skewness and kurtosis statistics can be evaluated visually via a histogram.  They can also be calculated by hand.  This is generally unnecessary with modern statistical software (such as Minitab)


Further Information:  Variance is a measure of the variation around the mean.  It measures how far a set of data points are spread out from their mean.  The units are the square of the units used for the original data. • For example, a variable measured in meters will have a variance measured in meters squared.  It is the square of the standard deviation. - Descriptive Statistics Variance = s



The Inter Quartile Range is also used to measure variability.  Quartiles divide an ordered data set into 4 parts.  Each contains 25% of the data.  The inter quartile range contains the middle 50% of the data (i.e. Q3-Q1).  It is often used when the data is not normally distributed.




 


评论

此博客中的热门博文

The Ultimate Tool Stack for AI Agents

弦线驻波