Skip to content
View ashlesha10's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report ashlesha10

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
ashlesha10/README.md

Header

Typing SVG

Portfolio LinkedIn GitHub

Hello, data world! 👋

I'm a Boston-based Senior Data Engineer with 8+ years of experience turning high-volume, messy, business-critical data into dependable platforms and useful decisions.

My favorite engineering problems live where scale, reliability, and business context meet. I have built and supported cloud data systems processing approximately 1.3 billion events every day, redesigned workflows to run 60% faster, and introduced controls that reduced production failures by approximately 70%.

ashlesha = {
    "focus": ["scalable pipelines", "trusted data", "analytics platforms"],
    "daily_tools": ["Python", "SQL", "Airflow", "AWS", "Redshift"],
    "engineering_style": "observable, maintainable, and business-aware",
    "currently_exploring": ["Databricks", "PySpark", "Agentic AI"]
}

My data-engineering fingerprint 🧬

1.3B 10+ 60% 70% 5 hrs
events/day sources unified faster processing fewer failures earlier reporting

From raw events to real decisions ⚡

flowchart LR
    A["10+ Data Sources"] --> B["Python + Airflow"]
    B --> C["Amazon S3"]
    C --> D["Redshift Models"]
    D --> E["Trusted Analytics"]
    E --> F["Product + Business Decisions"]
Loading

What I bring to that pipeline:

  • Ingestion: high-volume batch and near-real-time integrations
  • Orchestration: parallel execution, sensors, retries, backfills, and reusable DAG components
  • Modeling: dimensional models, reusable data marts, and semantic datasets
  • Reliability: validation, reconciliation, anomaly detection, schema-drift monitoring, and lineage
  • Consumption: product analytics, operational metrics, executive dashboards, and self-service BI

My toolbox 🧰

Python SQL Apache Airflow Apache Spark AWS Amazon S3 Amazon Redshift Databricks MongoDB Docker Power BI Tableau Grafana Git

Layer Technologies & practices
Build Python, SQL, PySpark, ETL/ELT, batch processing
Orchestrate Apache Airflow, dependency management, retries, backfills
Scale AWS, S3, Redshift, Glue, Athena, Spark, Databricks
Model Data warehouses, lakes, marts, dimensional modeling
Trust Data quality, anomaly detection, schema drift, lineage
Deliver Power BI, Looker, Tableau, Grafana, self-service analytics

Engineering stories I’m proud of 🚀

🏗️ Cloud data platform at billion-event scale
  • Built Python, SQL, and Airflow pipelines integrating more than 10 enterprise, product, financial, and customer sources
  • Modeled curated datasets in Amazon Redshift for adoption, engagement, ARR, churn, experimentation, and renewals
  • Delivered trusted operational metrics and analytics across large-scale product deployments
⚙️ Faster, calmer pipeline operations
  • Refactored DAGs with parallel execution, sensor-based dependencies, retries, backfills, and reusable components
  • Reduced end-to-end processing time by 60% and delivered executive reports five hours earlier
  • Added structured monitoring and validation, reducing production pipeline failures by approximately 70%
☁️ Enterprise warehouse migration
  • Helped migrate more than one billion payroll and financial records from legacy systems to a cloud data warehouse
  • Re-engineered ETL logic using Python and SQL
  • Used parallel runs, source-to-target testing, reconciliation, and production validation to protect reporting continuity
🔬 Scientific analytics automation
  • Created a Python pipeline to ingest, transform, validate, and standardize approximately 500 LC-MS datasets
  • Automated preparation and reporting so research teams could reach analytical results faster

What I’m building next 🌱

  • Production-style Databricks + PySpark pipelines
  • A data-quality framework with validation, observability, and alerting
  • Practical agentic AI tools for data engineering workflows
  • Technical write-ups about the difficult parts of running pipelines beyond the happy path

GitHub pulse 📊

Ashlesha's GitHub statistics Ashlesha's most used languages

Activity Graph

Beyond the pipeline 💡

I care about more than moving data from point A to point B. The best platforms help people trust the numbers, understand the context, and make better decisions. That is the kind of engineering work I want my repositories—and my career—to represent.

Have an interesting data problem? Let’s talk.

LinkedIn Portfolio

Footer

Pinned Loading

  1. Data-Analysis-and-Machine-Learning-Projects Data-Analysis-and-Machine-Learning-Projects Public

    Forked from rhiever/Data-Analysis-and-Machine-Learning-Projects

    Repository of teaching materials, code, and data for my data analysis and machine learning projects.

    Jupyter Notebook

  2. Logistic-Regression-in-Python-Project Logistic-Regression-in-Python-Project Public

    Forked from pb111/Logistic-Regression-in-Python-Project

    Logistic Regression in Python Project

    Jupyter Notebook

  3. Recruit-Restaurant-Visitor-Forecasting Recruit-Restaurant-Visitor-Forecasting Public

    Recruit Restaurant Visitor Forecasting

    Jupyter Notebook 1

  4. Women-s-E-Commerce-Clothing-Reviews Women-s-E-Commerce-Clothing-Reviews Public

    Predictive Analytics

    Jupyter Notebook 1 1

  5. Intermediate_Projects Intermediate_Projects Public

    R