I'm a Boston-based Senior Data Engineer with 8+ years of experience turning high-volume, messy, business-critical data into dependable platforms and useful decisions.
My favorite engineering problems live where scale, reliability, and business context meet. I have built and supported cloud data systems processing approximately 1.3 billion events every day, redesigned workflows to run 60% faster, and introduced controls that reduced production failures by approximately 70%.
ashlesha = {
"focus": ["scalable pipelines", "trusted data", "analytics platforms"],
"daily_tools": ["Python", "SQL", "Airflow", "AWS", "Redshift"],
"engineering_style": "observable, maintainable, and business-aware",
"currently_exploring": ["Databricks", "PySpark", "Agentic AI"]
}| 1.3B | 10+ | 60% | 70% | 5 hrs |
|---|---|---|---|---|
| events/day | sources unified | faster processing | fewer failures | earlier reporting |
flowchart LR
A["10+ Data Sources"] --> B["Python + Airflow"]
B --> C["Amazon S3"]
C --> D["Redshift Models"]
D --> E["Trusted Analytics"]
E --> F["Product + Business Decisions"]
What I bring to that pipeline:
- Ingestion: high-volume batch and near-real-time integrations
- Orchestration: parallel execution, sensors, retries, backfills, and reusable DAG components
- Modeling: dimensional models, reusable data marts, and semantic datasets
- Reliability: validation, reconciliation, anomaly detection, schema-drift monitoring, and lineage
- Consumption: product analytics, operational metrics, executive dashboards, and self-service BI
| Layer | Technologies & practices |
|---|---|
| Build | Python, SQL, PySpark, ETL/ELT, batch processing |
| Orchestrate | Apache Airflow, dependency management, retries, backfills |
| Scale | AWS, S3, Redshift, Glue, Athena, Spark, Databricks |
| Model | Data warehouses, lakes, marts, dimensional modeling |
| Trust | Data quality, anomaly detection, schema drift, lineage |
| Deliver | Power BI, Looker, Tableau, Grafana, self-service analytics |
🏗️ Cloud data platform at billion-event scale
- Built Python, SQL, and Airflow pipelines integrating more than 10 enterprise, product, financial, and customer sources
- Modeled curated datasets in Amazon Redshift for adoption, engagement, ARR, churn, experimentation, and renewals
- Delivered trusted operational metrics and analytics across large-scale product deployments
⚙️ Faster, calmer pipeline operations
- Refactored DAGs with parallel execution, sensor-based dependencies, retries, backfills, and reusable components
- Reduced end-to-end processing time by 60% and delivered executive reports five hours earlier
- Added structured monitoring and validation, reducing production pipeline failures by approximately 70%
☁️ Enterprise warehouse migration
- Helped migrate more than one billion payroll and financial records from legacy systems to a cloud data warehouse
- Re-engineered ETL logic using Python and SQL
- Used parallel runs, source-to-target testing, reconciliation, and production validation to protect reporting continuity
🔬 Scientific analytics automation
- Created a Python pipeline to ingest, transform, validate, and standardize approximately 500 LC-MS datasets
- Automated preparation and reporting so research teams could reach analytical results faster
- Production-style Databricks + PySpark pipelines
- A data-quality framework with validation, observability, and alerting
- Practical agentic AI tools for data engineering workflows
- Technical write-ups about the difficult parts of running pipelines beyond the happy path
I care about more than moving data from point A to point B. The best platforms help people trust the numbers, understand the context, and make better decisions. That is the kind of engineering work I want my repositories—and my career—to represent.