Data pipelines that survive production.
5+ years across banking and property tech, regularly trusted with architecture decisions and cross-team leadership. Batch and streaming, from raw source to analytics-ready — with CI/CD wired in.
- 124M+records / day
- 75%infra cost cut
- 99.8%pipeline reliability
- 35Ktxn / minute streaming
01 About
Data Engineer with 5+ years in banking and property tech, owning architecture decisions and leading delivery across teams. I design medallion-layered warehouses, real-time streaming pipelines, and orchestration that people can debug in minutes instead of hours.
Most of my work is measured in what it removed: cost, latency, manual steps, and 3 a.m. pages.
02 Experience
-
May 2026 — Present
Data Engineer
UNBLK Pte Ltd (Ohmyhome) · Singapore, remote
- Led orchestration migration from Cloud Composer to Windmill, cutting infra cost 75% (US$1,200 → US$300/mo) and consolidating scattered DAGs into one execution model.
- Fixed a broken home-valuation pipeline, restoring completeness from ~2 to 38 rows/day (~19x) and improving downstream valuation accuracy.
- Re-architected fragmented sources into a dbt-modeled BigQuery medallion warehouse (5 datasets), taking worst-case root-cause debugging from 8 hours to under 15 minutes.
- Rebuilt MATCH lead-matching lineage in 4 days: 113 → 69 stages, downstream consumers from a dozen+ to 4.
Windmill · dbt · BigQuery · Python · curl_cffi · GCP
-
Sep 2025 — Mar 2026
Data Engineer — Big Data, Streaming & Analytics
Bank Negara Indonesia · Jakarta
- Shipped a production Kafka + Spark Streaming pipeline handling 25–35K transactions/minute, with schema registry enforcing schema evolution and data contracts.
- Delivered a unified datamart for 20+ business and campaign reports, cutting query turnaround from a full day to under 5 minutes over 134M+ daily records.
- Built HC data architecture consolidating 17+ sources into a 5-layer medallion datamart, 10+ entity tables for AI-ready analytics on ~36K employees.
- Held production Spark workflows at 99.8% reliability while improving job performance up to 92%.
Kafka · Spark Streaming · PySpark · Hive/Impala · Schema Registry
-
May 2023 — Nov 2025
Data Engineer — Lead Automation & Campaign Analytics
Bank Negara Indonesia · Jakarta
- Automated the sales-lead ETL end to end (passworded Excel, FTP delivery, email alerts), taking processing from 2 hours to under 5 minutes.
- Optimized fuzzy matching on 2M+ records with RapidFuzz + multiprocessing: accuracy +30%, runtime −50%.
- Built a self-healing SQL retry layer — failure resolution time −80%, 100% query success without manual watching.
- Owned delivery and mentored junior and vendor engineers to shorten onboarding.
Python · PySpark · RapidFuzz · SQL · Bash
-
May 2021 — Apr 2023
Data Engineer — Data Marts & Query Optimization
Bank Negara Indonesia · Jakarta
- Automated daily/weekly/monthly datamart refreshes via CDSW — 100% on-time delivery, zero manual runs.
- Partitioned Hive/Impala tables on Parquet: 3x faster queries on less storage.
- Tuned slow SQL across Hadoop, cutting execution time up to 70% and serving 10+ business domains.
Hive · Impala · Parquet · PySpark · Cloudera
Before data: civil engineering & construction — quantity surveyor, structural/finishing inspector, drafter (2016–2020) · STEM teacher, part-time (2013–2014).
03 Projects
Banking Loan Risk ELT
How do banks screen millions of loan records daily? End-to-end pipeline from raw files to risk-ready tables.
Airflow · PySpark · dbt · BigQuery · Docker
E-Commerce ELT on Cloud Composer
Production-shaped Airflow 2.x pipeline with CI/CD and Slack alerting, runnable end to end.
Cloud Composer · dbt · BigQuery · GitHub Actions
Three-Source Sales ELT
Extracts from API, Postgres and Cloud Storage into one modeled sales summary.
Airflow · Docker · dbt · GCP
Banking Streaming OLTP → OLAP
Real-time transaction stream landed into analytics storage without breaking schema contracts.
Kafka · Spark Streaming · Schema Registry
Massive Lead Assignment
Distributes 100K–1M leads across sales reps by geography and customer criteria, fairly, round-robin.
Pandas · PySpark · Hive · CML
High-Speed Fuzzy Matching
Millions of messy name records matched in parallel instead of overnight.
RapidFuzz · Multiprocessing · Python
04 Skills
Languages
Python · SQL · Bash
Orchestration
Airflow · Windmill · dbt · Dagster · Cloud Composer · Cron
Streaming & Big Data
Kafka · Spark Streaming · PySpark · Hadoop · Hive · Impala · HDFS · Databricks
Cloud & Warehouse
BigQuery · Cloud Storage · Cloud Run · GKE · Vertex AI · Snowflake · Medallion architecture · Parquet · Partitioning
DevOps & CI/CD
Docker · Kubernetes · GitHub Actions · Bitbucket Pipelines · Linux · Git
Databases
PostgreSQL · MySQL · Oracle Exadata · Teradata · HiveQL · ImpalaQL
Scraping & Automation
curl_cffi (anti-bot) · Scrapy · BeautifulSoup · Playwright · Selenium · RapidFuzz
Education — Data Science & Machine Learning, Purwadhika Digital Technology School (2020) · B.Eng. Civil Engineering, Pembangunan Jaya University (2013–2017)
05 Contact
Open to data engineering work — streaming, warehouse design, orchestration rescue missions.