TA
Focus: LLM Applications - GenAI - Data Engineering - MLOps - Cloud Infrastructure - Engineering Leadership
AI & Data Engineer. I build LLM workflows for real operations and the data platforms behind them. I use Amazon Bedrock and Claude to turn unstructured advertising-material instructions into structured data. The workflow automates 10% of a critical operational process.
Before that, I engineered data systems that handled 6B records at peak, saved $7K/month in infrastructure costs, and grew an engineering team from 3 to 7 before splitting it into two teams.
AWS Certified Solutions Architect: Associate (July 2022).
Data engineering
- 6B records at peak
- $7K/month infrastructure savings: EMR $7K -> $3K, Azure reporting $1K/month, and self-hosted Elasticsearch about $2K/month
- 12 h -> <8 h complex ETL processing
- 60K -> 250K pages/hour crawler throughput (4.17x)
Leadership
- Grew an engineering team from 3 to 7, then split it into two teams; established KPIs and OKRs
See theanh.github.io for the long versions.
- DoodleCraft: child-focused sketch-to-crayon generation with Stable Diffusion 1.5 and a rank-16 LoRA trained on 53 images.
- SketchNet: convolutional neural network (CNN) for hand-drawn sketches. 95.08% test accuracy, 938 KB model, and 0.98 ms model-only CPU latency per sample. Live demo.
- DIA Risk Screener: five model families on 477 compounds. Model disagreement is an uncertainty signal; it is not a clinical decision tool. Live demo.
- PCA Audio Toolkit: Principal Component Analysis (PCA)-based audio denoising and lossy compression. Live demo.
- Cutting Spark shuffle cost: wide vs narrow transformations on billion-record EMR pipelines.
Python - TypeScript - JavaScript - PHP - Amazon Bedrock - Claude - PyTorch - scikit-learn - XGBoost - Apache Airflow - n8n - PySpark - Datadog - AWS - Elasticsearch - Docker - Terraform