Abstract
Automated data cleaning is becoming more and more important because of the rapidly increasing amounts of data. Manual control of data is very difficult and time consuming. Therefore one must know about the different anomalies that may occur in their data to be able to identify and clean occurring anomalies. The "Oberösterreichische Gebietskrankenkasse" is aware of that problem. Due to that fact it is important to implement a new tool for ETL process control and to handle the huge amount of data faster and more efficient.
This work summarizes data cleaning techniques and reviews how the term data quality is defined in current literature. Therefore a classification of data quality dimensions is made. Furthermore this work presents different occurrences of poor data quality and categorizes them. It also illustrates that data cleaning is an iterative process in all state-of-the-art data cleaning methods.
From the practical viewpoint, this work analyzed and summarizes the methods and processes implemented in five different data cleaning tools. Each tool was evaluated, how efficiently it supports the goals of the ETL process control planned by the OÖGKK. This work contains the specification of a framework for an automated data cleaning process aligned for the OÖGKK (SofaP). Computed or saved reference values and boundary values are used for analyzing and data cleaning. The framework provides modularity. Therefore adaptations can be made easily.
| Translated title of the contribution | Analysing Data Quality on Loading Data in Data Warehouses on the example of ETL-Processes of the OÖGKK |
|---|---|
| Original language | German (Austria) |
| Supervisors/Reviewers |
|
| Publication status | Published - Sept 2008 |
Fields of science
- 102 Computer Sciences
- 102015 Information systems
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver