DATA ENGINEER | MICROSOFT AZURE · PYTHON · SQL
I Engineer Data
That Moves.
5+ years' combined experience across data engineering and data testing, specialising in Microsoft Azure, Python, SQL and end-to-end ETL/ELT pipeline delivery — from raw ingestion to trusted, production-ready platforms.
Python
SQL
PySpark
Azure
Databricks
Fabric
From Raw Data → To Business Insight
API
→
Ingestion
→
Bronze
→
Silver
→
Gold
→
Warehouse
→
BI
My Data Engineering Journey
From diverse data sources to trusted insights — building the pipelines that power decisions.
☁
DATA SOURCES
APIs · Files · Databases
Events · Streaming
Events · Streaming
→
⚙
INGESTION
Azure Data Factory
Python · APIs
Python · APIs
→
✦
PROCESSING
Databricks · PySpark
Spark
Spark
→
🗄
DATA LAKE
ADLS Gen2 · Delta Lake
SCD1 / SCD2
SCD1 / SCD2
→
🏠
WAREHOUSE
Synapse · SQL
Microsoft Fabric
Microsoft Fabric
→
📊
CONSUMPTION
Power BI · Analytics
Applications
Applications
Technical Arsenal
The tools and technologies I use to build, manage and scale data platforms.
🐍
Programming
- Python (Pandas, PySpark)
- Requests, JSON, Pathlib
- Azure SDK
- SQL & T-SQL
△
Azure & Cloud
- Azure Data Factory
- Data Lake Storage
- Blob Storage
- Azure SQL Database
- Synapse Analytics
- Key Vault · Monitor
- Databricks · Azure DevOps
🗄
Data Engineering
- ETL / ELT
- Data Integration & Transformation
- Data Cleansing
- Data Validation & Quality
- Data Modelling & Warehousing
- Batch Processing
- API Integration
◈
Testing & QA
- Data & ETL Testing
- Source-to-Target Validation
- Data Reconciliation
- Regression & Functional Testing
- Defect Management
- Databricks Validation Frameworks
◎
Formats & Tools
- CSV, JSON, Parquet
- Git · Azure DevOps
- VS Code · Jupyter Notebook
- Power BI
Featured Project
Employee Data Modernisation — Scalable Data Platform
Employee Data Modernisation
Scalable Data Platform
Built an end-to-end data pipeline to ingest employee data from REST APIs, transform and model using Databricks, and load to a data warehouse for analytics.
?
ProblemInconsistent data, manual processes, no single source of truth.
⚙
SolutionAutomated pipeline with data validation, SCD1/SCD2 and orchestration.
↑
ImpactImproved data quality, reduced processing time and enabled real-time reporting.
REST API
→
Python / ADF
→
Azure Data Lake
→
Databricks
→
Bronze (Raw)
Silver SCD1 (Updates)
Silver SCD2 (History)
SQL / Fabric
→
Power BI
Experience Timeline
My journey in data, from analysis to engineering.
2017 – 2020
BSc Computer Science & Mathematics
Loyola
Foundations in programming, data structures & statistics
Jun 2020 – Nov 2021
Data Tester
London, UK
▸Source-to-target validation, SQL
▸Reconciliation · Defect management (Azure DevOps)
▸Reconciliation · Defect management (Azure DevOps)
Dec 2021 – Present
Data Engineer
London, UK
▸Azure Data Factory · Databricks · PySpark
▸ETL/ELT, data quality & pipeline monitoring
▸ETL/ELT, data quality & pipeline monitoring
Engineering Principles
How I build reliable, scalable and maintainable data systems.
🛡
Reliable
before fast
⚙
Automated
before manual
◎
Observable
before production
📈
Scalable
by design
✓
Data quality
is part of the pipeline
GitHub Activity
Consistent contributions to open source and personal projects.
24
REPOSITORIES
1,240+
COMMITS
12
FOLLOWERS
TOP LANGUAGES
Python
38%
SQL
22%
PySpark
16%
JavaScript
8%
Other
16%
Last 6 months
Let's Work Together
I'm open to new opportunities and would love to discuss how I can add value to your team.
🗓 Or book a meeting directly — calendly.com/data96561 ↗