0% found this document useful (0 votes)

42 views21 pages

03preprocessing Part1

The chapter discusses data preprocessing, which includes data cleaning, integration, reduction, and transformation. These steps are necessary to improve data quality and prepare it for data mining. Data cleaning involves handling incomplete, noisy, and inconsistent data through techniques such as filling in missing values, smoothing noisy data, and resolving inconsistencies. The goals of data cleaning are to ensure data quality dimensions such as accuracy, completeness, consistency and interpretability.

Uploaded by

baigsalman251

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PPTX, PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

42 views21 pages

03preprocessing Part1

Uploaded by

baigsalman251

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PPTX, PDF, TXT or read online on Scribd

You are on page 1/ 21

Data Mining

Dr. Shahid Mahmood Awan

http://turing.cs.pub.ro/mas_11
curs.cs.pub.ro
shahid.awan@umt.edu.pk
University of Management and Technology

Fall 2017
Data Mining:
Concepts and Techniques
(3rd ed.)

— Chapter 3 —

Jiawei Han, Micheline Kamber, and Jian Pei

University of Illinois at Urbana-Champaign &
Simon Fraser University
©2011 Han, Kamber & Pei. All rights reserved.
2
Chapter 3: Data Preprocessing

 Data Preprocessing: An Overview

 Data Quality
 Major Tasks in Data Preprocessing
 Data Cleaning
 Data Integration
 Data Reduction
 Data Transformation and Data Discretization
 Summary
3
Data Quality: Why Preprocess the Data?


Model Accuracy  Data Quality
 GIGO {garbage in, garbage out}

 Data Quality: accuracy, completeness,

consistency, timeliness, believability, and
interpretability.

4
Imagine that you are a manager at AllElectronics and have been
charged with analyzing the company’s data with respect to your
branch’s sales. You immediately set out to perform this task. You
carefully inspect the company’s database and data warehouse,
identifying and selecting the attributes or dimensions (e.g., item,
price, and units sold) to be included in your analysis. Alas! You notice
that several of the attributes for various tuples have no recorded
value. For your analysis, you would like to include information as to
whether each item purchased was advertised as on sale, yet you
discover that this information has not been recorded. Furthermore,
users of your database system have reported errors, unusual values,
and inconsistencies in the data recorded for some transactions. In
other words, the data you wish to analyze by data mining techniques
are incomplete (lacking attribute values or certain attributes of
interest, or containing only aggregate data); inaccurate or noisy
(containing errors, or values that deviate from the expected); and
inconsistent (e.g., containing discrepancies in the department codes
used to categorize items).Welcome to the real world!

5
Data Preprocessing
 This scenario illustrates three of the elements defining data quality:
accuracy, completeness, and consistency.

 Inaccurate, incomplete, and inconsistent data are

commonplace properties of large real-world databases and data
warehouses.

6
Data Quality: Why Preprocess the Data?

 Measures for data quality: A multidimensional view

 Accuracy: correct or wrong, accurate or not
 Completeness: not recorded, unavailable, …
 Consistency: some modified but some not, dangling, …
 Timeliness: timely update?
 Believability: how trustable the data are correct?
 Interpretability: how easily the data can be
understood?

7
Major Tasks in Data Preprocessing
 Data cleaning
 Fill in missing values, smooth noisy data, identify or remove
outliers, and resolve inconsistencies
 Data integration
 Integration of multiple databases, data cubes, or files
 Data reduction
 Dimensionality reduction
 Numerosity reduction
 Data compression
 Data transformation and data discretization
 Normalization
 Concept hierarchy generation

8
Major Tasks in Data Preprocessing

9
Knowledge Discovery Process
 Data mining: the core
of knowledge Knowledge Interpretation
discovery process.
Data Mining

Task-relevant Data
Data transformations

Preprocessed Selection
Data
Data Cleaning

Data Integration

Databases
Chapter 3: Data Preprocessing

 Data Preprocessing: An Overview

 Data Quality
 Major Tasks in Data Preprocessing
 Data Cleaning
 Data Integration
 Data Reduction
 Data Transformation and Data Discretization
 Summary
11
Data Cleaning
 Data in the Real World Is Dirty: Lots of potentially incorrect data,
e.g., instrument faulty, human or computer error, transmission error
 incomplete: lacking attribute values, lacking certain attributes of
interest, or containing only aggregate data
 e.g., Occupation=“ ” (missing data)
 noisy: containing noise, errors, or outliers
 e.g., Salary=“−10” (an error)
 inconsistent: containing discrepancies in codes or names, e.g.,
 Age=“42”, Birthday=“03/07/2010”
 Was rating “1, 2, 3”, now rating “A, B, C”
 discrepancy between duplicate records
 Intentional (e.g., disguised missing data)
 Jan. 1 as everyone’s birthday?
12
Incomplete (Missing) Data

 Data is not always available

 E.g., many tuples have no recorded value for several
attributes, such as customer income in sales data
 Missing data may be due to
 equipment malfunction
 inconsistent with other recorded data and thus
deleted
 data not entered due to misunderstanding
 certain data may not be considered important at the
time of entry
 not register history or changes of the data
 Missing data may need to be inferred 13
How to Handle Missing Data?
 Ignore the tuple: usually done when class label is missing
(when doing classification)—not effective when the % of
missing values per attribute varies considerably
 Fill in the missing value manually: tedious + infeasible?
 Fill in it automatically with
 a global constant : e.g., “unknown”, a new class?!
 the attribute mean
 the attribute mean for all samples belonging to the
same class: smarter
 the most probable value: inference-based such as
Bayesian formula or decision tree
14
15
16
Noisy Data
 Noise: random error or variance in a measured variable
 Incorrect attribute values may be due to
 faulty data collection instruments

 data entry problems

 data transmission problems

 technology limitation

 inconsistency in naming convention

 Other data problems which require data cleaning

 duplicate records

 incomplete data

 inconsistent data

17
How to Handle Noisy Data?
 Binning
 first sort data and partition into (equal-frequency) bins

 then one can smooth by bin means, smooth by bin

median, smooth by bin boundaries, etc.

 Regression
 smooth by fitting the data into regression functions

 Clustering
 detect and remove outliers

 Combined computer and human inspection

 detect suspicious values and check by human (e.g.,

deal with possible outliers)

18
19
20
Data Cleaning as a Process
 Data discrepancy detection
 Use metadata (e.g., domain, range, dependency, distribution)

 Check field overloading

 Check uniqueness rule, consecutive rule and null rule

 Use commercial tools

 Data scrubbing: use simple domain knowledge (e.g., postal

code, spell-check) to detect errors and make corrections

 Data auditing: by analyzing data to discover rules and

relationship to detect violators (e.g., correlation and clustering

to find outliers)
 Data migration and integration
 Data migration tools: allow transformations to be specified

 ETL (Extraction/Transformation/Loading) tools: allow users to

specify transformations through a graphical user interface

 Integration of the two processes
 Iterative and interactive (e.g., Potter’s Wheels)

Data Preprocessing Part 1
No ratings yet
Data Preprocessing Part 1
14 pages
03 Preprocessing
No ratings yet
03 Preprocessing
18 pages
Day-4 Preprocessing
No ratings yet
Day-4 Preprocessing
11 pages
VIPDMTheory Chapter 3
No ratings yet
VIPDMTheory Chapter 3
87 pages
Data Preprocessing - Cleaning and Normalization
No ratings yet
Data Preprocessing - Cleaning and Normalization
11 pages
Aiml Data Preprocessing
No ratings yet
Aiml Data Preprocessing
99 pages
DataPreprocessing 2
No ratings yet
DataPreprocessing 2
68 pages
Lecture Source: Books by Tan, Steinbach, Kumar Han, Kamber & Pei Evans Dinesh Kumar + Experiential Knowledge
No ratings yet
Lecture Source: Books by Tan, Steinbach, Kumar Han, Kamber & Pei Evans Dinesh Kumar + Experiential Knowledge
40 pages
Pre Processing
No ratings yet
Pre Processing
43 pages
Chapter 3& 4
No ratings yet
Chapter 3& 4
60 pages
DWDM LS3 Fall 24 25
No ratings yet
DWDM LS3 Fall 24 25
50 pages
DEC - Unit II Data Pre-Processing
No ratings yet
DEC - Unit II Data Pre-Processing
96 pages
Unit - II
No ratings yet
Unit - II
56 pages
Correlation
No ratings yet
Correlation
14 pages
Estimasi Anggaran Biaya Google Adwords Iklan Website
No ratings yet
Estimasi Anggaran Biaya Google Adwords Iklan Website
54 pages
3 Preprocessing
No ratings yet
3 Preprocessing
27 pages
Data Preprocessing 1 - Annotated
No ratings yet
Data Preprocessing 1 - Annotated
23 pages
Module2 DataPreprocessing
No ratings yet
Module2 DataPreprocessing
27 pages
Data Mining - Lecture 2
No ratings yet
Data Mining - Lecture 2
23 pages
DM Chapter 3
No ratings yet
DM Chapter 3
60 pages
Unit 1datapre Processing Datacleaningtransformationreductionintegration 240509092339 7095c9af
No ratings yet
Unit 1datapre Processing Datacleaningtransformationreductionintegration 240509092339 7095c9af
88 pages
03 Data Preprocessing
No ratings yet
03 Data Preprocessing
15 pages
02 Data - Preprocessing - 4,5,6
No ratings yet
02 Data - Preprocessing - 4,5,6
54 pages
Data Preparation Guide COS10022
No ratings yet
Data Preparation Guide COS10022
61 pages
Data Preprocessing Essentials
No ratings yet
Data Preprocessing Essentials
9 pages
Data Preprocessing
No ratings yet
Data Preprocessing
48 pages
Dmi Unit 3
No ratings yet
Dmi Unit 3
12 pages
Data Preprocessing
No ratings yet
Data Preprocessing
15 pages
Data Preprocessing Techniques Guide
No ratings yet
Data Preprocessing Techniques Guide
32 pages
DWM
No ratings yet
DWM
14 pages
18mca52c U2
No ratings yet
18mca52c U2
23 pages
Pre Processing
No ratings yet
Pre Processing
68 pages
DM Unit 3
No ratings yet
DM Unit 3
15 pages
Data Preprocessing Essentials
No ratings yet
Data Preprocessing Essentials
41 pages
Preprocessing
No ratings yet
Preprocessing
90 pages
Unit 2 Preprocessing
No ratings yet
Unit 2 Preprocessing
39 pages
Chapter 3: Data Preprocessing
No ratings yet
Chapter 3: Data Preprocessing
63 pages
Mod2 DM
No ratings yet
Mod2 DM
86 pages
3 DSEngineering
No ratings yet
3 DSEngineering
64 pages
03 Preprocessing
No ratings yet
03 Preprocessing
59 pages
Data Preprocessing
No ratings yet
Data Preprocessing
77 pages
Data Cleaning Preprocessing
No ratings yet
Data Cleaning Preprocessing
28 pages
Pre Processing
No ratings yet
Pre Processing
52 pages
Module 2 (C) - Data Preprocessing
No ratings yet
Module 2 (C) - Data Preprocessing
50 pages
UNIT - 2 .DataScience 04.09.18
No ratings yet
UNIT - 2 .DataScience 04.09.18
53 pages
Datapreparation
No ratings yet
Datapreparation
59 pages
Data Pre Processing
No ratings yet
Data Pre Processing
48 pages
Concepts and Techniques: - Chapter 3
No ratings yet
Concepts and Techniques: - Chapter 3
63 pages
04 DM BI Data Preprocessing
No ratings yet
04 DM BI Data Preprocessing
93 pages
Data Preprocessing Essentials
No ratings yet
Data Preprocessing Essentials
33 pages
DM Lect3
No ratings yet
DM Lect3
41 pages
Chapter 3: Data Preprocessing
No ratings yet
Chapter 3: Data Preprocessing
30 pages
Data Mining Requires Collecting Great Amount of Data (Available in Data Warehouses or Databases) To Achieve The Intended Objective
No ratings yet
Data Mining Requires Collecting Great Amount of Data (Available in Data Warehouses or Databases) To Achieve The Intended Objective
37 pages
Lecture 3 Unit 1
No ratings yet
Lecture 3 Unit 1
61 pages
Data Preprocessing for Tech Students
No ratings yet
Data Preprocessing for Tech Students
59 pages
Color Image Processing Guide
No ratings yet
Color Image Processing Guide
40 pages
Ch10-Image Segmentation
No ratings yet
Ch10-Image Segmentation
22 pages
Image Processing Techniques
No ratings yet
Image Processing Techniques
65 pages
Ch05-Image Restoration
No ratings yet
Ch05-Image Restoration
49 pages
03preprocessing3 Part3 4
No ratings yet
03preprocessing3 Part3 4
49 pages
02data Part4
No ratings yet
02data Part4
28 pages
02data Part2
No ratings yet
02data Part2
34 pages
01 Intro 1
No ratings yet
01 Intro 1
50 pages
02data Part1
No ratings yet
02data Part1
19 pages
Palombini - 1993 - Machine Songs V Pierre Schaeffer From Research I
No ratings yet
Palombini - 1993 - Machine Songs V Pierre Schaeffer From Research I
7 pages
EWB - Vehicle Tracking Mobile App
No ratings yet
EWB - Vehicle Tracking Mobile App
16 pages
Cab Tilt
100% (1)
Cab Tilt
39 pages
How Is Numerical Integration Programmed in S7-SCL and STEP 7
No ratings yet
How Is Numerical Integration Programmed in S7-SCL and STEP 7
2 pages
Mean Free Path and Collision Time
No ratings yet
Mean Free Path and Collision Time
2 pages
Low Cost Solar Simulators 2015
No ratings yet
Low Cost Solar Simulators 2015
2 pages
Self-Awareness and Regulation: Session 2
No ratings yet
Self-Awareness and Regulation: Session 2
73 pages
Quadratic Equation (Short Notes)
No ratings yet
Quadratic Equation (Short Notes)
3 pages
B.Tech Project Report Template
No ratings yet
B.Tech Project Report Template
15 pages
Act 7-Me Lab
No ratings yet
Act 7-Me Lab
7 pages
Nirmal Resume
No ratings yet
Nirmal Resume
3 pages
SIP-106 GHG Emissions Inventory For Asphalt Mix Production in The US - NAPA June 2022
No ratings yet
SIP-106 GHG Emissions Inventory For Asphalt Mix Production in The US - NAPA June 2022
34 pages
Numerical Methods for Engineering Students
No ratings yet
Numerical Methods for Engineering Students
2 pages
Macro-Fungi Diversity in Mt. Kitanglad
No ratings yet
Macro-Fungi Diversity in Mt. Kitanglad
72 pages
Chemistry Investigatory Project Class 12th
100% (1)
Chemistry Investigatory Project Class 12th
22 pages
Critical Thinking and Effective Writing Skills For The Professional Accountant
No ratings yet
Critical Thinking and Effective Writing Skills For The Professional Accountant
13 pages
04-Ceragon-IP-10G Radio Configuration PDF
No ratings yet
04-Ceragon-IP-10G Radio Configuration PDF
16 pages
American International University-Bangladesh (Aiub)
100% (3)
American International University-Bangladesh (Aiub)
9 pages
Is Science Dangerous
No ratings yet
Is Science Dangerous
1 page
Size Reduction in Crushing & Grinding
No ratings yet
Size Reduction in Crushing & Grinding
8 pages
Poshan Rijal Project Final
No ratings yet
Poshan Rijal Project Final
32 pages
Understanding Various Grief Models
No ratings yet
Understanding Various Grief Models
8 pages
Quotation LBM - Meet K Drama 2025
No ratings yet
Quotation LBM - Meet K Drama 2025
4 pages
CV - Viraj Sandaruwa Jul23
No ratings yet
CV - Viraj Sandaruwa Jul23
1 page
High-Speed SSDs for Gamers
No ratings yet
High-Speed SSDs for Gamers
1 page
Collaborative Feedback Form
No ratings yet
Collaborative Feedback Form
2 pages
Ai Use Cases Msme
No ratings yet
Ai Use Cases Msme
13 pages
Guidance For Generative AI in Education and Research
No ratings yet
Guidance For Generative AI in Education and Research
48 pages
Cessna 182 Skylane Certification
No ratings yet
Cessna 182 Skylane Certification
24 pages
Bank Exam Prep & Current Affairs
No ratings yet
Bank Exam Prep & Current Affairs
21 pages

03preprocessing Part1

Uploaded by

03preprocessing Part1

Uploaded by

Data Mining

Dr. Shahid Mahmood Awan

Jiawei Han, Micheline Kamber, and Jian Pei

 Data Preprocessing: An Overview

 Data Quality: accuracy, completeness,

 Inaccurate, incomplete, and inconsistent data are

 Measures for data quality: A multidimensional view

 Data Preprocessing: An Overview

 Data is not always available

 data entry problems

 data transmission problems

 inconsistency in naming convention

 Other data problems which require data cleaning

 then one can smooth by bin means, smooth by bin

median, smooth by bin boundaries, etc.

 Combined computer and human inspection

deal with possible outliers)

 Check field overloading

 Check uniqueness rule, consecutive rule and null rule

 Use commercial tools

 Data scrubbing: use simple domain knowledge (e.g., postal

code, spell-check) to detect errors and make corrections

relationship to detect violators (e.g., correlation and clustering

 ETL (Extraction/Transformation/Loading) tools: allow users to

specify transformations through a graphical user interface

You might also like