Veidlapa Nr. M-3 (8)
Study Course Description

Data Pre-processing

Main Study Course Information

Course Code
FK_083
Branch of Science
Other medical sciences; Other Sub-Branches of Medical Sciences
ECTS
5.00
Target Audience
Business Management; Health Management; Management Science
LQF
Level 7
Study Type And Form
Full-Time

Study Course Implementer

Course Supervisor
Structure Unit Manager
Structural Unit
Department of Physics
Contacts

Riga, Anniņmuižas boulevard 26a, 1st floor, office 147a and b, fizika@rsu.lv, +371 67061539

About Study Course

Objective

The objective of the data pre-processing study course is to provide students with essential skills, to prepare raw data for analysis. The main objectives include: Understanding data pre-processing: to understand the importance of data pre-processing and the basics of data analysis workflow. Data cleaning: to learn methods how to process missing values, remove duplicates, and correct errors to ensure data accuracy and consistency. Data transformation: to transform data into suitable formats for analysis, including normalisation, scaling and categorical variable encoding. Feature engineering: to create new features from existing data to improve model performance. Invalid data processing: to identify and manage invalid data to prevent deviations during analysis. Data integration and reduction: to combine data from different sources and reduce size for effective analysis. Practical experience: to obtain practical experience with real-world datasets using industry standard tools and software. Best practices and tools: to learn best practices and familiarise with tools and libraries such as Python’s Pandas, R and SQL. Preparation for improved analysis: to ensure readiness to perform additional data analysis tasks such as machine learning and statistical analysis. Ethical considerations: to discuss ethical aspects, including data privacy and security during pre-processing. At the end of the course, students will be able to convincingly prepare raw data for different analytical applications, ensuring that they are clean, well-structured and ready to use.

Preliminary Knowledge

Knowledge of informatics at secondary school level.

Learning Outcomes

Knowledge

1.After completing the “Data Pre-Processing” study course, students will gain in-depth knowledge of data pre-processing methods and techniques in various data formats and carriers, and understand the importance of data quality and its impact on data analysis.

Skills

1.During the study course, students will develop practical skills in importing, cleaning, transforming, and extracting features from various data sources and formats. They will be able to process missing values, detect anomalies, and address data imbalances.

Competences

1.Having completed the study course, students will be competent to perform a full cycle of data pre-processing across different projects, effectively addressing real-world problems, be able to adapt to different data types and processing challenges, develop automated solutions, and prepare data for further analysis and modeling. Students will be prepared to work in the fields of data science and analytics, applying acquired knowledge and skills in a professional environment.

Assessment

Individual work

Title
% from total grade
Grade
1.

Individual work

30.00% from total grade
Test

Students independently perform practical tasks and submit practical work reports in the e-learning environment.

Examination

Title
% from total grade
Grade
1.

Examination

70.00% from total grade
10 points

Develop a project in which students perform data preprocessing on a dataset. Present the project and evaluate the results.

Study Course Theme Plan

FULL-TIME
Part 1
  1. Lecture

Modality
Location
Contact hours
On site
Auditorium
2

Topics

Data Science Foundations (L1). PACE Strategy. Data preparation for analysis. Role of pre-processing in data analysis and machine learning processes. From raw data to finished data: key steps and methods. Data types and formatting.
  1. Class/Seminar

Modality
Location
Contact hours
On site
Auditorium
3

Topics

Basics of data preprocessing (P1). From raw data to prepared data. Initial data research and analysis. Evaluation of features and initial problems.
  1. Class/Seminar

Modality
Location
Contact hours
On site
Auditorium
3

Topics

Descriptive Statistics as Quality Control (P2). Fundamental concepts and measures of descriptive statistics. Principles of data description and summarization. Calculation and interpretation of measures of central tendency and dispersion. Identification of data distributions, anomalies, and deviations. Application of descriptive statistics in data quality assessment and quality control.
  1. Class/Seminar

Modality
Location
Contact hours
On site
Auditorium
3

Topics

Data Type Conversion and Category Harmonization (P4). Ensuring consistent data types and units of measurement within a dataset. Data normalization and scaling methods. Standardization and harmonization of categorical data across different data sources. Category encoding approaches and their application in statistical analysis and machine learning models. Principles for ensuring data comparability, consistency, and analytical quality.
  1. Class/Seminar

Modality
Location
Contact hours
On site
Auditorium
3

Topics

Creating New Groups, Filtering, and Working with Date Fields (P5). Creating new variables and risk groups using logical conditions and derived indicators. Principles of data filtering and selection using simple and advanced filters. Application of AND/OR logic in defining medical criteria.
Description

Working with date fields and harmonization of date formats. Calculation of age, time intervals, seasons, and other derived variables from date data. Quality assessment of date fields and their practical application in patient grouping and data analysis.

  1. Class/Seminar

Modality
Location
Contact hours
On site
Auditorium
3

Topics

Filtering Outliers and Erroneous Values and Merging Datasets (P6). Principles for identifying outliers and erroneous data. Use of logical thresholds, statistical criteria (Z-score, IQR), and clinical context in data quality assessment. Filtering, classification, and documentation of inappropriate and extreme values.
Description

Fundamental principles of dataset merging, use of common identifiers, and ensuring data consistency. Application of different types of joins for data integration, duplicate control, and quality assessment of the merged data.

  1. Class/Seminar

Modality
Location
Contact hours
On site
Auditorium
3

Topics

Missing Data, Duplicates, and Consistency (P3). Data organization skills. Data transformation and manipulation. Data filtering, selection, and grouping. Identification and imputation of missing data. Detection and removal of duplicate and inconsistent values. Methods for ensuring data quality.
  1. Class/Seminar

Modality
Location
Contact hours
On site
Auditorium
3

Topics

Data Reduction and Balancing (P7). Characteristics of high-dimensional data and dimensionality reduction methods. Feature Selection and Feature Extraction approaches and their application in statistical analysis and machine learning. Principles of data splitting for creating representative samples and training models.
Description

The need for data balancing in the presence of imbalanced classes. Data balancing methods, their advantages, and limitations in medical data analysis and predictive model development.

  1. Class/Seminar

Modality
Location
Contact hours
On site
Auditorium
3

Topics

Summary and Next Steps (P8). Review and systematization of the key stages of the data preprocessing process. Summary of data quality principles, common data-related issues, and approaches to addressing them. Data acquisition resources and publicly available databases relevant to the field. Introduction to statistical software, data visualization, and programming tools (Google Colab) for data processing.
  1. Class/Seminar

Modality
Location
Contact hours
On site
Auditorium
3

Topics

Practical Data Preprocessing (P9). Real-world applications. Data cleaning and transformation in real-world projects. Data preparation for obtaining final results.
  1. Class/Seminar

Modality
Location
Contact hours
On site
Auditorium
3

Topics

Handling Sensitive Data (P10). Types of sensitive data and their identification and classification in medical and research datasets. Principles of personal data protection, General Data Protection Regulation (GDPR) requirements, and ethical considerations when working with sensitive information. Fundamental principles of secure data storage, access control, and encryption. Pseudonymization and anonymization methods, their applications, and limitations. Solutions for protecting sensitive data thr
  1. Class/Seminar

Modality
Location
Contact hours
On site
Auditorium
3

Topics

Project (P11). Preprocessing of a dataset. Data cleaning, transformation, and preparation for data analysis. Project presentation, evaluation of results, and application of the acquired skills.
  1. Class/Seminar

Modality
Location
Contact hours
On site
Auditorium
3

Topics

Project (P12). Preprocessing of a dataset. Data cleaning, transformation, and preparation for data analysis. Project presentation, evaluation of results, and application of the acquired skills.
Total ECTS (Creditpoints):
5.00
Contact hours:
38 Academic Hours
Final Examination:
Exam

Bibliography

Required Reading

1.

Hands-On Data Preprocessing in Python. EBSCOhost Ebook Academic Collection, 2022.Suitable for English stream

2.

Data Wrangling with PythonSuitable for English stream

Additional Reading

1.

Foundational Python for Data ScienceSuitable for English stream

2.

Python for Data ScienceSuitable for English stream

Other Information Sources

1.

Preprocessing - Categorical DataSuitable for English stream

2.

PacktPublishing/Hands-On-Data-Preprocessing-in-PythonSuitable for English stream

3.

How to Preprocess Data in PythonSuitable for English stream