I’m a Data Engineer who enjoys turning messy data into clean, reliable datasets.
My sweet spot is data ingestion, transformations, and performance tuning I like making pipelines faster, cheaper, and easier to trust.
I work on:
- Building batch + streaming pipelines
- Handling multiple file formats (CSV, JSON, Parquet, Avro, etc.)
- Writing clean transformations (SQL / PySpark / Python)
- Improving performance (partitioning, joins, caching, pruning, cost tuning)
- Adding data quality checks and monitoring so failures get caught early
Right now I’m focused on:
- Strengthening my end-to-end pipeline projects (realistic datasets + production-style structure)
- Writing better tests + data quality validations
- Sharpening performance patterns in Spark + SQL
Languages: SQL, Python
Big Data: Spark / PySpark / Databricks
Orchestration: Airflow
Cloud: AWS / Azure / GCP
Warehouses: Snowflake / BigQuery / Redshift
Dev: Git, GitHub Actions, Docker
Data Quality/Observability: Great Expectations / Deequ / CloudWatch / Datadog
When I’m not building pipelines:
- 🏐 Beach volleyball
- 📷 Photography
- 🚁 FPV drones (landscape flying & filming)
- 🍜 Trying new cuisines
- 🎧 Classic + hip-hop
