Building a Medallion Architecture on Databricks (AWS)
The Medallion Architecture has become the de facto standard for organizing data in a lakehouse. In this post, we expl...
Databricks • dbt • Airflow • PySpark • SQL
Staff-level Data Engineer with 11+ years building and owning large-scale data pipelines across four global financial-data firms. I work across Databricks, dbt, Airflow, PySpark and SQL — designing platform architecture, treating data quality as a first-class concern, and making governed data genuinely self-serve for analytics and data-science teams.
This site shares architecture patterns, technical decisions and personal projects — deliberately keeping my employer's project specifics and internal volumes private.
The practices I own as a Staff Data & AI Engineer
Own ingestion and transformation end-to-end on Databricks, conforming heterogeneous sources into a single trusted schema — fault-tolerant PySpark where schemas demand it, config-driven dbt models where they don't.
Validation gates as a first-class concern — dbt tests, DLT expectations and Great Expectations — plus quarantine flows for bad records and end-to-end lineage, so issues are traceable instead of argued about.
Treat platform cost as an engineering outcome, not a bill. Workload profiling and autoscaling tuning across environments, with delivery SLAs held steady while spend comes down.
Standardized deployment with Databricks Asset Bundles and GitHub Actions, so pipelines ship through a consistent, reviewable dev → prod path instead of relying on tribal knowledge.
Reference patterns I build and own — architecture and reasoning, without employer-specific detail
A medallion lakehouse built for clarity and governance — Bronze, Silver and Gold layers on Databricks with Delta Lake and Unity Catalog to keep sources trustable as they evolve.
A governed Gold layer surfaced via OneLake and Direct Lake in Microsoft Fabric — delivering near-real-time BI to stakeholders without duplicating the curated data.
Config-driven transformations and orchestrated DAGs — the separation I apply in practice between runtime-schema-driven PySpark and repeatable dbt models, all tested through parity checks.
Validation gates, quarantine flows and lineage — applying dbt tests, DLT expectations and Great Expectations so pipelines fail fast and issues are traceable instead of opaque.
Eleven years owning data platforms end-to-end across four global financial-data firms
Scopes below are described in general terms — volumes, team sizes and internal project names are intentionally omitted.
My own builds — GitHub repos I publish to show how I work
Aggregated and normalized real-time vendor pricing for US structured finance securities into one consistent store for the pricing desk.
View repository →Conformed multi-vendor market data into analysis-ready, point-in-time feature tables serving analysts and data scientists.
View repository →Built the ingestion, feature-engineering and train/serve-parity pipeline feeding downstream error-prediction and root-cause models.
View repository →First cloud build: a deployed data-quality tool for yield-curve validation, provisioned on AWS (EC2/S3/Redshift) with Infrastructure as Code.
View repository →Orchestrated an Azure data platform with a zoned lake and star schema, serving always-current self-service BI to stakeholders.
View repository →My own build: a medallion architecture for large-scale international-trade analytics on Databricks — the lakehouse patterns I work with daily, reproduced end-to-end in public.
View repository →End-to-end automated integration with data quality as a first-class concern — validation gates, quarantine and a reusable ingestion template.
View repository →My exploration of retrieval-augmented generation over financial research — prototyping grounded, cited retrieval on Databricks Mosaic AI Vector Search, the reference pattern behind my move into Data & AI Engineering.
View repository →A unified Fabric lakehouse serving BI (Direct Lake), analysts (SQL) and data science from one governed Gold layer — showing the same architecture delivered on a second cloud platform.
View repository →Deep dives into data pipeline architecture, Databricks, dbt, Airflow, and CI/CD for data
The Medallion Architecture has become the de facto standard for organizing data in a lakehouse. In this post, we expl...
Empowering the Telugu tech community — enterprise data engineering, demystified.
తెలుగులో డేటా ఇంజనీరింగ్ — మన భాషలో, మన కోసం.
డేటాబ్రిక్స్ లేక్హౌస్ — మొదటి నుండి ప్రొడక్షన్ వరకు
Your complete roadmap to mastering Databricks — clusters, notebooks, Delta Lake, Unity Catalog, and production-grade pipelines. Let's build this together!
మైక్రోసాఫ్ట్ ఫ్యాబ్రిక్ — యూనిఫైడ్ అనలిటిక్స్ విప్లవం
OneLake, Lakehouses, Data Factory, and Direct Lake mode — everything you need to architect modern analytics in Fabric. This changes the game.
మెడాలియన్ ఆర్కిటెక్చర్ — బ్రాంజ్, సిల్వర్ & గోల్డ్ వివరణ
The architecture pattern powering modern lakehouses. I'll walk you through real-world implementations with Delta Live Tables on Databricks.
డేటా ఇంజనీర్ల కోసం పైస్పార్క్ — ఇంటర్వ్యూ & అంతకు మించి
Not just interview prep — real production patterns. Transformations, window functions, performance tuning, and the questions top companies actually ask.
యూనిటీ క్యాటలాగ్ — ఎంటర్ప్రైజ్ డేటా గవర్నెన్స్
Access control, data lineage, and quality enforcement at scale. I'll show you how to set up governance that actually works across multi-cloud Databricks.
డెల్టా లైవ్ టేబుల్స్ — డిక్లరేటివ్ ETL పైప్లైన్స్
Stop writing boilerplate. DLT lets you declare your pipeline logic and handles orchestration, quality, and recovery. Let me show you how the pros do it.
The modern data stack I build, own and operate on every day
Building dependable systems out of ambiguous requirements
Staff Data Engineer — Modern Data Stack (Databricks • dbt • Airflow)
I'm a Staff-level Data Engineer with 11+ years building and owning large-scale data pipelines across four global financial-data firms — Moody's, FactSet, Franklin Templeton and S&P Global. My focus areas are data pipeline architecture, distributed data processing, orchestration, data quality & observability, and CI/CD for data.
At Moody's I own ingestion and transformation for a Databricks platform at enterprise scale. I've led the migration of legacy PySpark pipelines onto a standardized dbt + PySpark stack orchestrated by Airflow, built the team's data-quality and observability framework, and treated platform cost as an engineering outcome rather than a bill to be paid.
I'm at my best in the grey zone between platform and practice — setting standards a team actually reuses, reviewing code, and mentoring engineers. I collaborate closely with analytics and data-science colleagues to make governed data genuinely self-serve, and recently delivered a retrieval-augmented generation proof-of-concept as a reference pattern for the team's future AI/ML work.
I keep this site deliberately general. It shares architecture patterns, technical reasoning and personal projects rather than my employer's project specifics — the depth is here, the confidential detail is not. The projects below are my own builds, and they're the most direct way to see how I work.
Beyond work, I create Telugu-language tutorials on YouTube, making data engineering concepts accessible to the Telugu-speaking tech community worldwide.
Engineering foundations, plus a finance lens on the business
Open to Staff Data Engineer roles in Data & AI Engineering