Data Pre-processing
Study Course Implementer
Riga, Anniņmuižas boulevard 26a, 1st floor, office 147a and b, fizika@rsu.lv, +371 67061539
About Study Course
Objective
The objective of the data pre-processing study course is to provide students with essential skills, to prepare raw data for analysis. The main objectives include: Understanding data pre-processing: to understand the importance of data pre-processing and the basics of data analysis workflow. Data cleaning: to learn methods how to process missing values, remove duplicates, and correct errors to ensure data accuracy and consistency. Data transformation: to transform data into suitable formats for analysis, including normalisation, scaling and categorical variable encoding. Feature engineering: to create new features from existing data to improve model performance. Invalid data processing: to identify and manage invalid data to prevent deviations during analysis. Data integration and reduction: to combine data from different sources and reduce size for effective analysis. Practical experience: to obtain practical experience with real-world datasets using industry standard tools and software. Best practices and tools: to learn best practices and familiarise with tools and libraries such as Python’s Pandas, R and SQL. Preparation for improved analysis: to ensure readiness to perform additional data analysis tasks such as machine learning and statistical analysis. Ethical considerations: to discuss ethical aspects, including data privacy and security during pre-processing. At the end of the course, students will be able to convincingly prepare raw data for different analytical applications, ensuring that they are clean, well-structured and ready to use.
Preliminary Knowledge
Knowledge of informatics at secondary school level.
Learning Outcomes
Knowledge
1.After completing the “Data Pre-Processing” study course, students will gain in-depth knowledge of data pre-processing methods and techniques in various data formats and carriers, and understand the importance of data quality and its impact on data analysis.
Skills
1.During the study course, students will develop practical skills in importing, cleaning, transforming, and extracting features from various data sources and formats. They will be able to process missing values, detect anomalies, and address data imbalances.
Competences
1.Having completed the study course, students will be competent to perform a full cycle of data pre-processing across different projects, effectively addressing real-world problems, be able to adapt to different data types and processing challenges, develop automated solutions, and prepare data for further analysis and modeling. Students will be prepared to work in the fields of data science and analytics, applying acquired knowledge and skills in a professional environment.
Assessment
Individual work
|
Title
|
% from total grade
|
Grade
|
|---|---|---|
|
1.
Individual work |
30.00% from total grade
|
Test
|
|
Students independently perform practical tasks and submit practical work reports in the e-learning environment. |
||
Examination
|
Title
|
% from total grade
|
Grade
|
|---|---|---|
|
1.
Examination |
70.00% from total grade
|
10 points
|
|
Develop a project in which students perform data preprocessing on a dataset. Present the project and evaluate the results. |
||
Study Course Theme Plan
-
Lecture
|
Modality
|
Location
|
Contact hours
|
|---|---|---|
|
On site
|
Auditorium
|
2
|
Topics
|
Data Science Foundations (L1). PACE Strategy. Data preparation for analysis. Role of pre-processing in data analysis and machine learning processes. From raw data to finished data: key steps and methods. Data types and formatting.
|
-
Class/Seminar
|
Modality
|
Location
|
Contact hours
|
|---|---|---|
|
On site
|
Auditorium
|
3
|
Topics
|
Basics of data preprocessing (P1). From raw data to prepared data. Initial data research and analysis. Evaluation of features and initial problems.
|
-
Class/Seminar
|
Modality
|
Location
|
Contact hours
|
|---|---|---|
|
On site
|
Auditorium
|
3
|
Topics
|
Descriptive Statistics as Quality Control (P2). Fundamental concepts and measures of descriptive statistics. Principles of data description and summarization. Calculation and interpretation of measures of central tendency and dispersion. Identification of data distributions, anomalies, and deviations. Application of descriptive statistics in data quality assessment and quality control.
|
-
Class/Seminar
|
Modality
|
Location
|
Contact hours
|
|---|---|---|
|
On site
|
Auditorium
|
3
|
Topics
|
Data Type Conversion and Category Harmonization (P4). Ensuring consistent data types and units of measurement within a dataset. Data normalization and scaling methods. Standardization and harmonization of categorical data across different data sources. Category encoding approaches and their application in statistical analysis and machine learning models. Principles for ensuring data comparability, consistency, and analytical quality.
|
-
Class/Seminar
|
Modality
|
Location
|
Contact hours
|
|---|---|---|
|
On site
|
Auditorium
|
3
|
Topics
|
Creating New Groups, Filtering, and Working with Date Fields (P5). Creating new variables and risk groups using logical conditions and derived indicators. Principles of data filtering and selection using simple and advanced filters. Application of AND/OR logic in defining medical criteria.
Description
Working with date fields and harmonization of date formats. Calculation of age, time intervals, seasons, and other derived variables from date data. Quality assessment of date fields and their practical application in patient grouping and data analysis. |
-
Class/Seminar
|
Modality
|
Location
|
Contact hours
|
|---|---|---|
|
On site
|
Auditorium
|
3
|
Topics
|
Filtering Outliers and Erroneous Values and Merging Datasets (P6). Principles for identifying outliers and erroneous data. Use of logical thresholds, statistical criteria (Z-score, IQR), and clinical context in data quality assessment. Filtering, classification, and documentation of inappropriate and extreme values.
Description
Fundamental principles of dataset merging, use of common identifiers, and ensuring data consistency. Application of different types of joins for data integration, duplicate control, and quality assessment of the merged data. |
-
Class/Seminar
|
Modality
|
Location
|
Contact hours
|
|---|---|---|
|
On site
|
Auditorium
|
3
|
Topics
|
Missing Data, Duplicates, and Consistency (P3). Data organization skills. Data transformation and manipulation. Data filtering, selection, and grouping. Identification and imputation of missing data. Detection and removal of duplicate and inconsistent values. Methods for ensuring data quality.
|
-
Class/Seminar
|
Modality
|
Location
|
Contact hours
|
|---|---|---|
|
On site
|
Auditorium
|
3
|
Topics
|
Data Reduction and Balancing (P7). Characteristics of high-dimensional data and dimensionality reduction methods. Feature Selection and Feature Extraction approaches and their application in statistical analysis and machine learning. Principles of data splitting for creating representative samples and training models.
Description
The need for data balancing in the presence of imbalanced classes. Data balancing methods, their advantages, and limitations in medical data analysis and predictive model development. |
-
Class/Seminar
|
Modality
|
Location
|
Contact hours
|
|---|---|---|
|
On site
|
Auditorium
|
3
|
Topics
|
Summary and Next Steps (P8). Review and systematization of the key stages of the data preprocessing process. Summary of data quality principles, common data-related issues, and approaches to addressing them. Data acquisition resources and publicly available databases relevant to the field. Introduction to statistical software, data visualization, and programming tools (Google Colab) for data processing.
|
-
Class/Seminar
|
Modality
|
Location
|
Contact hours
|
|---|---|---|
|
On site
|
Auditorium
|
3
|
Topics
|
Practical Data Preprocessing (P9). Real-world applications. Data cleaning and transformation in real-world projects. Data preparation for obtaining final results.
|
-
Class/Seminar
|
Modality
|
Location
|
Contact hours
|
|---|---|---|
|
On site
|
Auditorium
|
3
|
Topics
|
Handling Sensitive Data (P10). Types of sensitive data and their identification and classification in medical and research datasets. Principles of personal data protection, General Data Protection Regulation (GDPR) requirements, and ethical considerations when working with sensitive information. Fundamental principles of secure data storage, access control, and encryption. Pseudonymization and anonymization methods, their applications, and limitations. Solutions for protecting sensitive data thr
|
-
Class/Seminar
|
Modality
|
Location
|
Contact hours
|
|---|---|---|
|
On site
|
Auditorium
|
3
|
Topics
|
Project (P11). Preprocessing of a dataset. Data cleaning, transformation, and preparation for data analysis. Project presentation, evaluation of results, and application of the acquired skills.
|
-
Class/Seminar
|
Modality
|
Location
|
Contact hours
|
|---|---|---|
|
On site
|
Auditorium
|
3
|
Topics
|
Project (P12). Preprocessing of a dataset. Data cleaning, transformation, and preparation for data analysis. Project presentation, evaluation of results, and application of the acquired skills.
|
Bibliography
Required Reading
Hands-On Data Preprocessing in Python. EBSCOhost Ebook Academic Collection, 2022.Suitable for English stream