Data Engineer — AI & Analytics

Jack Toke

I build reliable data pipelines & AI solutions that turn messy operational data into decisions.

Fig.01 — How I work: end-to-end data flow Batch + streaming
01
SOURCES
Calls · Ops · APIs
02
INGEST
Airflow · Beam · Pub/Sub
03
TRANSFORM
dbt · PySpark
04
WAREHOUSE
BigQuery · Delta
05
SERVE
Dash · Tableau
Databricks CertifiedQUT Grad Cert — Data AnalyticsLe WagonRMITBilingual Karen–English
01 / About

About

Based in
Ballarat region, Victoria
Focus
Data engineering + applied AI
Languages
English · Karen (bilingual)

I turn messy operational data into clean, trusted pipelines — and increasingly, into AI systems that act on it.

Data and IT professional with a Graduate Certificate in Data Analytics (QUT) and the Databricks Certified Data Engineer Associate credential. Hands-on across enterprise application support, systems integration, and business reporting.

At AME Systems I integrated Siemens Capital with enterprise systems and automated data synchronisation in Python & SQL. Now, working independently, I build AI and data engineering solutions for regional businesses — comfortable with stakeholders at every technical level.

02 / Skills & stack

Skills & tech stack

Data engineering
  • SQL
  • Python
  • Spark / PySpark
  • Databricks
  • dbt
  • Airflow
  • Apache Beam
  • Pub/Sub
Warehousing & ingestion
  • BigQuery
  • Postgres
  • Delta Lake
  • Auto Loader
  • Fivetran
  • Airbyte
Reporting & viz
  • Tableau
  • Plotly Dash
  • Streamlit
Platform & other
  • Docker & Compose
  • Unity Catalog
  • .NET (MAUI / ASP.NET Core)
  • Kotlin / Jetpack Compose
03 / Experience

Experience

Jan 2026 — Present Current

AI / Data Engineer · Freelance

AI telephone receptionist and lead-qualification, AI chatbots, automation and live reporting for regional SMBs.

Nov 2022 — Feb 2025

ICT Application Officer · AME Systems

Integrated Siemens Capital with enterprise systems (webhooks / SOAP); automated Python + SQL data synchronisation; built custom Sales and Engineering reporting; shipped a .NET MAUI + ASP.NET Core label app and Kotlin / Jetpack Compose Zebra scanner apps.

Oct 2020 — May 2022

Computer Operator · Hazeldene's Chicken Farm

Ran and maintained computerised packaging and label machinery; kept production and breakdown records.

May 2015 — Jun 2019

School Coordinator · Education Opportunity Foundation

Rolled out an English Teacher Toolkit across 20 schools on the Thai–Burmese border; trained 76 teachers; ran payroll for 20 schools.

04 / Projects

Selected work

More coming soon
01

Transport Victoria — GTFS-Realtime Pipeline

Production lakehouse measuring V/Line train punctuality from Public Transport Victoria’s GTFS feeds. A timer-triggered Azure Function (Python) downloads real-time protobuf trip updates to ADLS Gen2. A parallel Databricks job ingests the weekly static GTFS schedule across 11 bronze tables. Protobuf is decoded on the JVM with from_protobuf(), then modelled through a bronze–silver–gold medallion architecture with an SCD2 stops dimension. The gold fact flags on-time, severely late and skipped stops per served stop. Deployed serverless on Databricks with Unity Catalog and Bicep infrastructure-as-code, and served through a live Streamlit dashboard.

Azure Functions · Databricks · Protobuf · Medallion · Streamlit
02

Australian real estate lakehouse & analytics dashboard

End-to-end Databricks lakehouse over Australian residential property listings, ingesting buy, rent and sold data from the realty-in-au API. Raw JSON lands via Auto Loader into a bronze–silver–gold medallion architecture on Delta Lake and Azure Data Lake Storage (ADLS Gen2). dbt transforms it into a dimensional star schema that separates physical properties from listings, with conformed dimensions, bridge tables and snapshot facts. Hash-based surrogate keys, referential-integrity checks and dbt data-quality tests (uniqueness, not-null, relationships) keep the data trustworthy. The property-versus-listing split enables repeat-sales analysis across 2007–2026, identifying 640 properties sold twice and 82 sold three times. Orchestrated with decoupled Databricks Jobs for nightly ingestion and file-arrival-triggered transformation, deployed as infrastructure-as-code with Databricks Asset Bundles across separate dev and prod catalogs, and served through an interactive Streamlit dashboard.

Databricks · dbt · Auto Loader · Asset Bundles · CI/CD · Streamlit
03

Citi Bike Medallion Pipeline

Event-driven lakehouse over roughly 10 years of Citi Bike trip data (2016–present). Auto Loader ingests raw CSVs and canonicalises their changing historical schemas into one format in testable pure Python. Change Data Feed streams typed, deduplicated trips into silver via MERGE upserts. The gold layer recomputes only the affected days per batch, staying correct under late-arriving data and deletes. Deterministic ride IDs are derived by SHA-256 hashing for records that lack native identifiers. Deployed on serverless compute with three GitHub Actions workflows for CI, gated deployment and job runs, using service principals across separate dev, test and prod workspaces. Served through a live Streamlit dashboard.

Databricks · Auto Loader · CDC · Asset Bundles · CI/CD · Streamlit
04

End-to-end ELT — Airflow, dbt & DuckDB

Dockerised ELT pipeline over JSON receipt data from four small businesses, running fully self-contained with no external services. Apache Airflow DAGs orchestrate ingestion and trigger downstream transformation. Modular dbt models transform the data in DuckDB. A Streamlit dashboard answers business questions on seasonal profitability, customer loyalty and employee turnover.

Airflow · dbt · DuckDB · Streamlit · Docker
05

AI Telephone Receptionist & Lead Qualifier

Production AI voice system that answers inbound calls, qualifies and captures leads, and surfaces activity through live reporting for regional SMBs. Handles call routing, structured lead capture and automated follow-up, giving small businesses a first, practical step into AI adoption.

AI · LLM · Automation · Live reporting · Private repo
05 / Credentials

Certifications & education

Certifications
Databricks Certified Data Engineer Associate
Databricks · Lakehouse, Delta, Spark SQL & PySpark ELT, Auto Loader, Unity Catalog
Jul 2026
Data Engineering Bootcamp
Le Wagon · Airflow, dbt, PySpark, Pub/Sub, Beam, BigQuery
Jun 2025
Data Engineering
RMIT · Airflow + dbt, PySpark, Pub/Sub, Beam, BigQuery
Apr 2025
Education
Master of Biostatistics
The University of Adelaide
In progress
Graduate Certificate of IT Practice (Data Analytics)
Queensland University of Technology
Dec 2025
Web Development Bootcamp
Coder Academy
Feb 2020
BSc (Computer Science)
Victoria University
2010
06 / Contact

Get in touch

Let's build something reliable.

Open to data engineering roles and freelance projects. Drop a line and I'll get back to you.