• Home
  • Project
    Introduction
    • Introduction
    • Background and Motivation
    • What is a Good Timetable?
    • Project Aims and Scope
  • Graph Data
    Model
    • Graph vs Relational Data Models
    • Graph Data Model for Timetabling
    • Early Insights
    • Model Expansion
    • Graphing Time
  • Data
    Pipeline
    • ETL Overview
    • Approach
    • Configuration and Logging
    • Extract
    • Transform
    • Google Drive Load
    • Neo4j Load
    • Reflection
  • Timetable
    Metrics
    • Timetable Metrics
    • Metric Aggregations
    • Implementing Metrics
    • TQI Summary
  • Final
    Thoughts
  • Appendices
    & Extras
    • Appendix Table of Contents
    • References
    • Acknowledgements
  • Word
  1. Data Pipeline
  2. Transform
  • Home
  • Project Introduction
    • Introduction
    • Background and Motivation
    • What is a Good Timetable?
    • Project Aims and Scope
  • Graph Data Model
    • Graph vs Relational Data Models
    • Graph Data Model for Timetabling
    • Early Insights
    • Model Expansion
    • Graphing Time
  • Data Pipeline
    • ETL Overview
    • Approach
    • Configuration and Logging
    • Extract
    • Transform
    • Google Drive Load
    • Neo4j Load
    • Reflection
  • Timetable Metrics
    • Timetable Metrics
    • Metric Aggregations
    • Implementing Metrics
    • TQI Summary
  • Final Thoughts
  • Appendices
    • Random Graph Generator
    • Technology Stack
    • Configuration
    • Anonymisation
    • ETL Summary and Code
      • ETL Summary
      • ETL Code
      • Config and Misc
      • Extract-SQL
      • Extract
      • Google Drive Load
      • Transform
      • Neo4j Load
    • Neo4j & Cypher Code
      • Cypher Queries
      • Creating Nodes and Relationships
      • Deleting Nodes and Relationships
      • General Queries
      • Count Queries
      • Hard (timetabling) Constraints
      • Student Clashes
      • Soft Constraints
      • Rooms and Spaces
      • Perspectives
      • Blue Skies Opportunities
  • Supervision
    • Supervision
    • Notes Example 1
    • Notes Example 2
    • Notes Example 3
  • References
  • Acknowledgements

On this page

  • All data
  • Nodes and relationships
  1. Data Pipeline
  2. Transform

Transformation

TRANSFORM picks up where EXTRACT finished by using the extracted csv files as the source.

transform source_files CSV Files (./{hostkeys}/extract) validate_data Validating Data source_files->validate_data config Configuration config->validate_data clean_data Cleaning Data config->clean_data add_department Adding 'Department' to Nodes config->add_department anonymise_data π—”π—‘π—’π—‘π—¬π— π—œπ—¦π—œπ—‘π—š Personal Data config->anonymise_data augment_rooms Augmenting Rooms with Archibus Data config->augment_rooms create_relationships Creating Relationship Tables config->create_relationships validate_data->clean_data clean_data->add_department add_department->anonymise_data anonymise_data->augment_rooms augment_rooms->create_relationships processed_files Processed CSV Files create_relationships->processed_files

Configuration allows the user to control which nodes and relationships are included and how they are processed. There are options to specify validation, cleaning, data linking, anonymisation and relationship details.

It is also possible to specify datatypes. Neo4j assumes string datatype unless it is well-formatted or pre-determined. Config allows the user to specify specific datatypes like dates, times, point, Boolean, etc.

All data

  1. Validation - basic validation of the data is performed. Validation is extensible and can be expanded, as requirements are identified.
  2. Cleaned - basic cleaning of all data is performed by stripping empty space and removing non-printable characters, etc. using regex. The cleaning functionality is expandable.

With clean data, the transformation proper starts:

Nodes and relationships

  1. Add Organisational Unit - where appropriate, the University Organisational Unit (e.g. College, School, Department) is added to the node. This will be picked up as a property during load.
  2. Data Augmentation - Room data is augmented with additional properties from the location master database, including latitude, longitude, square meterage, etc. Data augmentation is extensible.
  3. Anonymisation - Personal data is anonymised. An anonymisation function was developed to remove and replace any personally identifiable information (PII). The pipeline extracts minimal PII but this is safely anonymised. The functional also adds fake emails. See Appendix for additional details
  4. Relationships - Based on requirements in the configuration, relationships are extracted including optional relationship properties.
Extract
Google Drive Load

Copyright 2024, Petter LΓΆvehagen

 

Built with Quarto