Staff Data Engineer • Data & AI Engineering

Modern Data Stack
at Enterprise Scale

Databricks • dbt • Airflow • PySpark • SQL

Staff-level Data Engineer with 11+ years building and owning large-scale data pipelines across four global financial-data firms. I work across Databricks, dbt, Airflow, PySpark and SQL — designing platform architecture, treating data quality as a first-class concern, and making governed data genuinely self-serve for analytics and data-science teams.

This site shares architecture patterns, technical decisions and personal projects — deliberately keeping my employer's project specifics and internal volumes private.

What I Bring

The practices I own as a Staff Data & AI Engineer

Pipeline Architecture

Own ingestion and transformation end-to-end on Databricks, conforming heterogeneous sources into a single trusted schema — fault-tolerant PySpark where schemas demand it, config-driven dbt models where they don't.

Data Quality & Observability

Validation gates as a first-class concern — dbt tests, DLT expectations and Great Expectations — plus quarantine flows for bad records and end-to-end lineage, so issues are traceable instead of argued about.

Cost Discipline

Treat platform cost as an engineering outcome, not a bill. Workload profiling and autoscaling tuning across environments, with delivery SLAs held steady while spend comes down.

CI/CD for Data

Standardized deployment with Databricks Asset Bundles and GitHub Actions, so pipelines ship through a consistent, reviewable dev → prod path instead of relying on tribal knowledge.

Professional Path

Eleven years owning data platforms end-to-end across four global financial-data firms

Scopes below are described in general terms — volumes, team sizes and internal project names are intentionally omitted.

Jan 2022 – Present • Bangalore, India

Assistant Director, Data Engineer Staff-level IC

Moody's

  • Own ingestion and transformation end-to-end for a large financial data platform on Databricks, conforming many heterogeneous global sources into a single trusted schema.
  • Led the migration of legacy PySpark pipelines onto a standardized dbt + PySpark stack orchestrated by Airflow, with automated value-parity checks comparing old and new outputs.
  • Built the team's data-quality and observability framework — validation gates, quarantine flows and Unity Catalog lineage — and set CI/CD for data with Databricks Asset Bundles and GitHub Actions.
  • Made the design call to separate runtime-schema-driven jobs (PySpark) from config-driven models (dbt) — a classification standard that keeps transformation logic testable and maintainable.
  • Published documented, event-driven data outputs and self-serve datasets consumed by several downstream analytics and data-science teams, and extended a governed Gold layer to Microsoft Fabric (Direct Lake) for near-real-time BI without duplication.
  • Provide technical direction and mentorship across the team, reviewing code to hold data-modeling and pipeline standards, and delivered the team's first GenAI proof-of-concept (RAG on Mosaic AI Vector Search) as a reference pattern for future AI/ML work.
Feb 2019 – Jan 2022 • Hyderabad, India

Senior Data Engineer Product Specialist II

FactSet Research Systems

  • Built high-throughput market-data pipelines on Apache Spark and AWS (EC2, S3, Redshift), provisioned via Terraform, delivering reliable pricing data to client-facing products.
  • Designed and shipped a data-quality web app applying unsupervised outlier detection to spot-rate data, catching pricing anomalies before client delivery.
  • Automated analytics workflows over vendor APIs, eliminating substantial recurring manual effort and turning ad-hoc product requests into self-serve pipelines.
  • Partnered with the product team to build repeatable usage-data pipelines, reducing delivery turnaround time.
Feb 2017 – Feb 2019 • Hyderabad, India

Data Engineer Research Analyst

Franklin Templeton Investments

  • Built ETL frameworks and data-warehousing solutions powering investment analytics and portfolio-management systems.
  • Designed dimensional data marts (star schema) giving research analysts and portfolio managers consistent, point-in-time datasets.
  • Automated recurring research workflows in Python and SQL, improving data freshness and freeing analyst time for higher-value work.
Oct 2014 – Jan 2017 • Hyderabad, India

Data Engineer Data Researcher II

S&P Global (via Genpact)

  • Built Python and SQL pipelines aggregating and normalizing multi-vendor pricing for US structured-finance securities into a single source for the pricing desk.
  • Automated structured-products operations and vendor feeds, improving processing turnaround time.
  • Established early data-quality and reconciliation checks that improved pricing accuracy and downstream trust.

Project Journey

My own builds — GitHub repos I publish to show how I work

2014

Structured Finance Pricing Pipeline

Python 2.7 · SQL — Data Engineer

Aggregated and normalized real-time vendor pricing for US structured finance securities into one consistent store for the pricing desk.

View repository →
2016

Market Data Pipeline & Feature Platform

Python · pandas — Data Engineer

Conformed multi-vendor market data into analysis-ready, point-in-time feature tables serving analysts and data scientists.

View repository →
2018

Error Telemetry Pipeline (DE for ML)

Python · NLP — Data Engineer → Senior Data Engineer

Built the ingestion, feature-engineering and train/serve-parity pipeline feeding downstream error-prediction and root-cause models.

View repository →
2019

Yield Curve Outlier Detection

AWS · Streamlit · Terraform — Senior Data Engineer

First cloud build: a deployed data-quality tool for yield-curve validation, provisioned on AWS (EC2/S3/Redshift) with Infrastructure as Code.

View repository →
2021

Platform Usage Analytics

Azure (ADF · Synapse · ADLS Gen2) · Power BI — Senior Data Engineer

Orchestrated an Azure data platform with a zoned lake and star schema, serving always-current self-service BI to stakeholders.

View repository →
2022

Customs & Trade Analytics Lakehouse

Databricks · PySpark · Delta · Unity Catalog — Assistant Director, Data Engineer (Staff-level IC)

My own build: a medallion architecture for large-scale international-trade analytics on Databricks — the lakehouse patterns I work with daily, reproduced end-to-end in public.

View repository →
2023

Grant Data Integration Pipeline

Databricks · Delta · Great Expectations — Assistant Director, Data Engineer (Staff-level IC)

End-to-end automated integration with data quality as a first-class concern — validation gates, quarantine and a reusable ingestion template.

View repository →
2025

Financial Research RAG

Databricks GenAI · Mosaic AI Vector Search · MLflow — Assistant Director, Data Engineer (Staff-level IC)

My exploration of retrieval-augmented generation over financial research — prototyping grounded, cited retrieval on Databricks Mosaic AI Vector Search, the reference pattern behind my move into Data & AI Engineering.

View repository →
2026 · Present

Enterprise Lakehouse on Microsoft Fabric

Microsoft Fabric · OneLake · Direct Lake — Assistant Director, Data Engineer (Staff-level IC)

A unified Fabric lakehouse serving BI (Direct Lake), analysts (SQL) and data science from one governed Gold layer — showing the same architecture delivered on a second cloud platform.

View repository →

Technical Blog

Deep dives into data pipeline architecture, Databricks, dbt, Airflow, and CI/CD for data

Data Architect Telugu

Empowering the Telugu tech community — enterprise data engineering, demystified.
తెలుగులో డేటా ఇంజనీరింగ్ — మన భాషలో, మన కోసం.

Series
తెలుగు · Telugu

Databricks Lakehouse — Zero to Production

డేటాబ్రిక్స్ లేక్‌హౌస్ — మొదటి నుండి ప్రొడక్షన్ వరకు

Your complete roadmap to mastering Databricks — clusters, notebooks, Delta Lake, Unity Catalog, and production-grade pipelines. Let's build this together!

Series
తెలుగు · Telugu

Microsoft Fabric — The Unified Analytics Revolution

మైక్రోసాఫ్ట్ ఫ్యాబ్రిక్ — యూనిఫైడ్ అనలిటిక్స్ విప్లవం

OneLake, Lakehouses, Data Factory, and Direct Lake mode — everything you need to architect modern analytics in Fabric. This changes the game.

Deep Dive
తెలుగు · Telugu

Medallion Architecture — Bronze, Silver & Gold Explained

మెడాలియన్ ఆర్కిటెక్చర్ — బ్రాంజ్, సిల్వర్ & గోల్డ్ వివరణ

The architecture pattern powering modern lakehouses. I'll walk you through real-world implementations with Delta Live Tables on Databricks.

Masterclass
తెలుగు · Telugu

PySpark for Data Engineers — Interview & Beyond

డేటా ఇంజనీర్ల కోసం పైస్పార్క్ — ఇంటర్వ్యూ & అంతకు మించి

Not just interview prep — real production patterns. Transformations, window functions, performance tuning, and the questions top companies actually ask.

Tutorial
తెలుగు · Telugu

Unity Catalog — Enterprise Data Governance

యూనిటీ క్యాటలాగ్ — ఎంటర్‌ప్రైజ్ డేటా గవర్నెన్స్

Access control, data lineage, and quality enforcement at scale. I'll show you how to set up governance that actually works across multi-cloud Databricks.

Hands-On
తెలుగు · Telugu

Delta Live Tables — Declarative ETL Pipelines

డెల్టా లైవ్ టేబుల్స్ — డిక్లరేటివ్ ETL పైప్‌లైన్స్

Stop writing boilerplate. DLT lets you declare your pipeline logic and handles orchestration, quality, and recovery. Let me show you how the pros do it.

Subscribe to Data Architect Telugu

Tech Stack

The modern data stack I build, own and operate on every day

Languages & Query

Python SQL PySpark

Data Platform

Databricks Unity Catalog Delta Lake Databricks Workflows Databricks Asset Bundles dbt Apache Spark

Orchestration

Apache Airflow MWAA Databricks Workflows Dependency-aware DAGs SLA alerting Backfills & retries

Data Quality & Observability

dbt tests DLT expectations Great Expectations Data reconciliation Unity Catalog lineage Quarantine flows

CI/CD for Data

Databricks Asset Bundles GitHub Actions Terraform (IaC)

Cloud

AWS S3 Glue MWAA Lambda IAM Redshift Microsoft Azure Microsoft Fabric OneLake Direct Lake

Data Modeling

Dimensional (star schema) Canonical modeling Medallion (Bronze/Silver/Gold) Data contracts

Data & AI Engineering

Mosaic AI Vector Search RAG patterns Isolation Forest Apache Kafka

About Me

Building dependable systems out of ambiguous requirements

Kamalakar Peta

Kamalakar Peta

Staff Data Engineer — Modern Data Stack (Databricks • dbt • Airflow)

I'm a Staff-level Data Engineer with 11+ years building and owning large-scale data pipelines across four global financial-data firms — Moody's, FactSet, Franklin Templeton and S&P Global. My focus areas are data pipeline architecture, distributed data processing, orchestration, data quality & observability, and CI/CD for data.

At Moody's I own ingestion and transformation for a Databricks platform at enterprise scale. I've led the migration of legacy PySpark pipelines onto a standardized dbt + PySpark stack orchestrated by Airflow, built the team's data-quality and observability framework, and treated platform cost as an engineering outcome rather than a bill to be paid.

I'm at my best in the grey zone between platform and practice — setting standards a team actually reuses, reviewing code, and mentoring engineers. I collaborate closely with analytics and data-science colleagues to make governed data genuinely self-serve, and recently delivered a retrieval-augmented generation proof-of-concept as a reference pattern for the team's future AI/ML work.

I keep this site deliberately general. It shares architecture patterns, technical reasoning and personal projects rather than my employer's project specifics — the depth is here, the confidential detail is not. The projects below are my own builds, and they're the most direct way to see how I work.

Beyond work, I create Telugu-language tutorials on YouTube, making data engineering concepts accessible to the Telugu-speaking tech community worldwide.

11+
Years Experience
4
Global Financial-Data Firms
3
Clouds Shipped On

Education

Engineering foundations, plus a finance lens on the business

2007 – 2011

B-Tech, Computer Science & Engineering

JNTU Anantapur

2012 – 2014

MBA, Finance

Sri Venkateswara University

Get in Touch

Open to Staff Data Engineer roles in Data & AI Engineering