Hello,

I'm Zain Shafique

Data engineer and analyst who builds end-to-end pipelines on Azure — from ADF and Databricks lakehouses to Power BI dashboards — with a focus on clean, reliable, decision-ready data.

About Me .

ZS

Data Engineer & Analyst

Islamabad, Pakistan

Software Engineering graduate from NUST who found a home in the data world. I design and build end-to-end data platforms on Azure: Medallion-architecture lakehouses, CDC pipelines, star schemas, and the Power BI dashboards that sit on top.

Alongside engineering, I manage accounting and financial reporting for a multi-location restaurant client and volunteered as a data analyst on a research project for Women in Sport, so I care about the decision the data drives, not just the pipeline that moves it.

3

Medallion architecture projects

6

Azure services used

~4.4K

Survey respondents analyzed

7

Production ADF pipeline patterns

Technical Skills

Python (Pandas, NumPy, Scikit-learn, TensorFlow)90%
SQL / T-SQL90%
Azure Data Factory90%
Azure Databricks / PySpark85%
Power BI / DAX85%
Azure Synapse & SQL Server80%
ETL/ELT, CDC & Data Modeling90%
Git / GitHub85%

Experience & Education .

Experience

01/2024 — Present

Freelance Financial Operations Specialist

US Restaurant Franchise Client · Remote · USA

  • Manage accounting and financial reporting for 4 restaurant locations using QuickBooks, including bookkeeping, reconciliations, and insurance administration
  • Prepare monthly P&L and sales reports for ownership across all locations
09/2025 — 04/2026

Data Analyst

Women in Sport × Statistics Without Borders · Remote · USA

  • Analyzed survey data from ~4, 400 youth respondents using Python, and SPSS, using cross - tabulations and visualizations to identify participation barriers by gender and ethnicity
  • Co-authored a team report on safety, confidence, and health - related barriers(e.g., 45.7 % of Asian females did not feel safe exercising outdoors), included as an appendix to the final report delivered to Women in Sport
06/2025 — 08/2025

AI Engineer Intern

Verior · Remote · Pakistan

  • Built a Flask microservice using the Cohere API for text generation and summarization, with a modular LLM wrapper, deployed on Vercel
  • Built a Node.js/Express translation API supporting 20 languages, with a JavaScript front end, deployed on Vercel
06/2024 — 09/2024

Data Science Intern

Byewise Limited · Remote · Pakistan

  • Built ETL workflows for data preparation and implemented sentiment analysis and anomaly detection models in Python

Education

08/2019 — 06/2023

B.E. in Software Engineering

National University of Sciences & Technology (NUST) · Islamabad, Pakistan

Skill Groups

Languages & Libraries

Python (Pandas, NumPy, Scikit-learn, TensorFlow)SQLT-SQLPySpark

Cloud & Data Engineering

Azure Data FactoryAzure DatabricksAzure Synapse AnalyticsADLS Gen2Azure SQL DBAzure Logic AppsGreat ExpectationsSQL ServerGit/GitHub

Data Modeling & BI

Medallion ArchitectureKimball Star SchemaSCD Type 1 & 2ETL/ELTCDCPower BIDAXREST APIs

Featured Projects .

ADF Pipeline Patterns — CDC, REST API & File Routing

  • 7 production-grade Azure Data Factory patterns: CDC incremental load with JSON & DB-based audit logging, dual REST API pagination, dynamic file routing and schema mapping
  • Delta Lake upserts via Mapping Data Flow and multi-table orchestration driven by a SQL control table with parallel execution
Azure Data FactoryDelta LakeCDCREST API

End-to-End Azure Data Lakehouse — ADF to Power BI

  • Full Medallion Architecture (Bronze → Silver → Gold) across Azure Data Factory, Databricks, and Synapse Analytics Serverless SQL on Adventure Works data
  • Delta Lake storage, parameterized Databricks notebooks with secret scopes, Great Expectations quality checks, and CI/CD via GitHub Actions
  • Analytics-ready data surfaced through Power BI, with T-SQL views on Synapse OPENROWSET for schema-on-read querying
AzureDatabricksSynapsePower BICI/CD

Databricks Medallion Pipeline

  • End-to-end lakehouse pipeline (Bronze → Silver → Gold) over e-commerce sales data using PySpark and Delta Lake
  • Incremental watermark loading; Gold layer modeled as a Kimball-style star schema with SCD Type 1 & 2 for full historical tracking
DatabricksPySparkDelta LakeSCD

Medallion Data Warehouse — SQL Server

  • End-to-end data warehouse on SQL Server implementing Medallion Architecture, integrating ERP and CRM source data via T-SQL ETL pipelines
  • Star schema (fact & dimension tables) in the Gold layer with data cleansing and standardization in Silver for optimized analytical queries
SQL ServerT-SQLStar SchemaETL