Data Mining

 



Data mining is one of the most useful techniques that help entrepreneurs, researchers, and individuals to etract valuable information from huge sets of data. Data mining is also called Knowledge Discovery in Database (KDD). The knowledge discovery process includes Data cleaning, Data integration, Data selection, Data transformation, Data mining, Pattern evaluation, and Knowledge presentation.

The process of etracting information to identify patterns, trends, and useful data that would allow the business to take the data-driven decision from huge sets of data is called Data Mining.


In other words, we can say that Data Mining is the process of investigating hidden patterns of information to various perspectives for categorization into useful data, which is collected and assembled in particular areas such as data warehouses, efficient analysis, data mining algorithm, helping decision making and other data requirement to eventually cost-cutting and generating revenue.

Data Mining is a process used by organizations to etract specific data from huge databases to solve business problems. It primarily turns raw data into useful information.


In the contet of computer science, “Data Mining” can be referred to as knowledge mining from data, knowledge etraction, data/pattern analysis, data archaeology, and data dredging.  It is basically the process carried out for the etraction of useful information from a bulk of data or data warehouses. 

Basically, Data mining has been integrated with many other techniques from other domains such as statistics, machine learning, pattern recognition, database and data warehouse systems, information retrieval, visualization, etc. to gather more information about the data and to helps predict hidden patterns, future trends, and behaviors and allows businesses to make decisions.


 Technically, data mining is the computational process of analyzing data from different perspectives, dimensions, angles and categorizing/summarizing it into meaningful information. 

Data Mining can be applied to any type of data e.g. Data Warehouses, Transactional Databases, Relational Databases, Multimedia Databases, Spatial Databases, Time-series Databases, World Wide Web. 


The whole process of Data Mining consists of three main phases: 


Data Pre-processing – Data cleaning, integration, selection, and transformation takes place

Data Etraction – Occurrence of eact data mining

Data Evaluation and Presentation – Analyzing and presenting results

You belong to the network analytical team of Mastercard. 


There is a high end restaurant in Toronto, called Shake & Fries. There is problem happening currently with them. The restaurant used to have a huge number of Loyal Mastercard customers. However, either those customers have stopped visiting, or the number of loyalists visiting has reduced drastically.


Since you have access to all the transactional data for all customers, they want to know from you:


What is going wrong currently? What might be affecting the footfall? Mention any specific metrics or methodologies that be used for recommendations?


 In this phase, business and data-mining goals are established.


First, you need to understand business and client objectives. You need to define what your client wants (which many times even they do not know themselves)

Take stock of the current data mining scenario. Factor in resources, assumption, constraints, and other significant factors into your assessment.

Using business objectives and current scenario, define your data mining goals.

A good data mining plan is very detailed and should be developed to accomplish both business and data mining goals.


In this phase, sanity check on data is performed to check whether its appropriate for the data mining goals.


First, data is collected from multiple data sources available in the organization.

These data sources may include multiple databases, flat filer or data cubes. There are issues like object matching and schema integration which can arise during Data Integration process. It is a quite comple and tricky process as data from various sources unlikely to match easily. For eample, table A contains an entity named cust_no whereas another table B contains an entity named cust-id.

Therefore, it is quite difficult to ensure that both of these given objects refer to the same value or not. Here, Metadata should be used to reduce errors in the data integration process.

Net, the step is to search for properties of acquired data. A good way to eplore the data is to answer the data mining questions (decided in business phase) using the query, reporting, and visualization tools.

Based on the results of query, the data quality should be ascertained. Missing data if any should be acquired.

In this phase, data is made production ready.


The data preparation process consumes about 90% of the time of the project.


The data from different sources should be selected, cleaned, transformed, formatted, anonymized, and constructed (if required).Data cleaning is a process to “clean” the data by smoothing noisy data and filling in missing values.


For eample, for a customer demographics profile, age data is missing. The data is incomplete and should be filled. In some cases, there could be data outliers. For instance, age has a value 300. Data could be inconsistent. For instance, name of the customer is different in different tables.


Data transformation operations change the data to make it useful in data mining. Following transformation can be applied

Data transformation operations would contribute toward the success of the mining process.


Smoothing: It helps to remove noise from the data.


Aggregation: Summary or aggregation operations are applied to the data. I.e., the weekly sales data is aggregated to calculate the monthly and yearly total.


Generalization: In this step, Low-level data is replaced by higher-level concepts with the help of concept hierarchies. For eample, the city is replaced by the county.


Normalization: Normalization performed when the attribute data are scaled up o scaled down. Eample: Data should fall in the range -2.0 to 2.0 post-normalization.


Attribute construction: these attributes are constructed and included the given set of attributes helpful for data mining.


The result of this process is a final data set that can be used in modeling.

In this phase, mathematical models are used to determine data patterns.


Based on the business objectives, suitable modeling techniques should be selected for the prepared dataset.

Create a scenario to test check the quality and validity of the model.

Run the model on the prepared dataset.

Results should be assessed by all stakeholders to make sure that model can meet data mining objectives.

In this phase, patterns identified are evaluated against the business objectives.


Results generated by the data mining model should be evaluated against the business objectives.


Gaining business understanding is an iterative process. In fact, while understanding, new business requirements may be raised because of data mining.


A go or no-go decision is taken to move the model in the deployment phase.

In the deployment phase, you ship your data mining discoveries to everyday business operations.


The knowledge or information discovered during data mining process should be made easy to understand for non-technical stakeholders.

A detailed deployment plan, for shipping, maintenance, and monitoring of data mining discoveries is created.

A final project report is created with lessons learned and key eperiences during the project. This helps to improve the organization’s business policy.

A relational database is a collection of multiple data sets formally organized by tables, records, and columns from which data can be accessed in various ways without having to recognize the database tables. Tables convey and share information, which facilitates data searchability, reporting, and organization.

A Data Warehouse is the technology that collects the data from various sources within the organization to provide meaningful business insights. The huge amount of data comes from multiple places such as Marketing and Finance. The etracted data is utilized for analytical purposes and helps in decision- making for a business organization. The data warehouse is designed for the analysis of data rather than transaction processing.

A combination of an object-oriented database model and relational database model is called an object-relational model. It supports Classes, Objects, Inheritance, etc.


One of the primary objectives of the Object-relational data model is to close the gap between the Relational database and the object-oriented model practices frequently utilized in many programming languages, for eample, C++, Java, C#, and so on.

A transactional database refers to a database management system (DBMS) that has the potential to undo a database transaction if it is not performed appropriately. Even though this was a unique capability a very long while back, today, most of the relational database systems support transactional database activities.

To create models, marketing companies use data mining. This was based on history to forecast who will respond to new marketing campaigns such as direct mail, online marketing, etc. This means that marketers can sell profitable products to targeted customers.

Since data etraction provides financial institutions information on loans and credit reports, data can determine good or bad credits by creating a model for historical customers. It also helps banks detect fraudulent transactions by credit cards that protect a credit card owner.

Data mining can motivate researchers to accelerate when the method analysis the data. Therefore they can work more time on other projects. Shopping behaviors can be detected. Most of the time, you may eperience new problems while designing specific shopping patterns. Therefore data mining is used to solve these problems. Mining methods can find all the information on these shopping patterns. This process also creates an area where all the unepected shopping patterns are calculated. This data etraction can be beneficial when shopping patterns are identified.

In marketing campaigns, mining techniques are used. This is to understand their own customers ‘ needs and habits. And from that, customers can also choose their brand’s clothes. Thus, you can definitely be self-reliant with the help of this technique. However, it provides possible information when it comes to decisions.

People use these data mining techniques to help them make some decisions in marketing or business. Today, with the use of this technology, all information can be determined. Also, using such technology, one can decide precisely what is unknown and unepected.

All information factors are part of the working nature of the system. The data mining systems can also be obtained from these. They can help you predict future trends, and with the help of this technology, this is entirely possible. And people also adopt behavioral changes.

We use data mining to find all kinds of unseen element information. And adding data mining helps you to optimize your website. Similarly, this data mining provides information that may use the technology of data mining.

With regards to insights, regression is the connection between an independent variable and a dependent variable in the statistical analysis methods. 

The line utilized in regression analysis charts and graphs means whether the connections between the factors are solid or frail, notwithstanding showing patterns throughout a particular measure of time. 

A true eperiment is a type of eperimental design and is thought to be the most accurate type of eperimental research. This is because a true eperiment supports or refutes a hypothesis using statistical analysis. A true eperiment is also thought to be the only eperimental design that can establish cause and effect relationships. So, what makes a true eperiment?


There are three criteria that must be met in a true eperiment

 Control group and eperimental group

 Researcher-manipulated variable

 Random assignment

Classification analysis is a data analysis task within data-mining, that identifies and assigns categories to a collection of data to allow for more accurate analysis. 


The classification method makes use of mathematical techniques such as decision trees, linear programming, neural network and statistics.


Classification analysis can be used to question, make a decision, or predict behavior through the use of an algorithm. It works by developing a set of training data which contains a certain set of attributes as well as the likely outcome. The job of the classification algorithm is to discover how that set of attributes reaches its conclusion.

Classification is a category of what is called supervised machine learning methods in which the data is split on two parts: the training set and the validation set. 


Using the training set, a model is learned by etracting the most discriminative features, which are already associated to know outputs.

There are two steps in the construction of a classification model.


Learning Step – this is where different algorithms are used to build a classifier by making the model learn using the training set available. The model has to be trained for the prediction of accurate results.


Classification Step - this is where the model used to predict class labels, tests the constructed model on test data. Which in turn estimates the accuracy of the classification rules.

Decision trees are among the most popular machine learning algorithms given their intelligibility and simplicity.

Decision tree builds classification or regression models in the form of a tree structure. 

It breaks down a dataset into smaller and smaller subsets while at the same time an associated decision tree is incrementally developed. 

The final result is a tree with decision nodes and leaf nodes. 

The goal is to create a model that predicts the value of a target variable based on several input variables.


In decision analysis, a decision tree can be used to visually and eplicitly represent decisions and decision making. 


In data mining, decision trees can be described also as the combination of mathematical and computational techniques to aid the description, categorization and generalization of a given set of data.

There are two main types of Decision Trees:

Classification Trees.

Regression Trees.


Classification trees (Yes/No types) :


What we’ve seen above is an eample of classification tree, where the outcome was a variable like ‘fit’ or ‘unfit’. Here the decision variable is Categorical/ discrete.


Such a tree is built through a process known as binary recursive partitioning. This is an iterative process of splitting the data into partitions, and then splitting it up further on each of the branches.


In this method a set of training eamples is broken down into smaller and smaller subsets while at the same time an associated decision tree get incrementally developed. At the end of the learning process, a decision tree covering the training set is returned.


The key idea is to use a decision tree to partition the data space into cluster (or dense) regions and empty (or sparse) regions.


In Decision Tree Classification a new eample is classified by submitting it to a series of tests that determine the class label of the eample. These tests are organized in a hierarchical structure called a decision tree. Decision Trees follow Divide-and-Conquer Algorithm.


Inepensive to construct.

Etremely fast at classifying unknown records.

Easy to interpret for small-sized trees

Accuracy comparable to other classification techniques for many simple data sets.

Ecludes unimportant features.

Easy to overfit.

Decision Boundary restricted to being parallel to attribute aes.

Decision tree models are often biased toward splits on features having a large number of levels.

Small changes in the training data can result in large changes to decision logic.

Large trees can be difficult to interpret and the decisions they make may seem counter intuitive.

Biomedical Engineering (decision trees for identifying features to be used in implantable devices).

Financial analysis (Customer Satisfaction with a product or service).

Astronomy (classify galaies).

System Control.

Manufacturing and Production (Quality control, Semiconductor manufacturing, etc).

Medicines (diagnosis, cardiology, psychiatry).

Physics (Particle detection).

评论

此博客中的热门博文

The Ultimate Tool Stack for AI Agents

弦线驻波