Solutions Architect and AI Systems Engineer
I turn fragile AI prototypes into observable, testable, production systems.
I have spent 14+ years designing backend platforms, enterprise workflows, distributed services, and production AI systems. My current work sits at the intersection of RAG quality, agent orchestration, event-driven architecture, and operational reliability.
- Building production AI architectures with explicit evaluation, security, and observability boundaries
- Designing Python/FastAPI services, event streams, idempotent workers, and reliable data workflows
- Turning system-design decisions into runnable reference implementations, tests, and operational documentation
- Contributing fixes upstream when the problem belongs in the ecosystem rather than a local workaround
Provider-independent FastAPI reference architecture for measurable retrieval-augmented generation.
- Hybrid BM25 and dense retrieval with explicit reranking
- Grounded citations and deterministic value-presence checks
- Reproducible evaluation with Recall@5 and MRR quality gates in CI
- Local adapters for offline development and provider interfaces for production integration
Repository | Architecture case study
Runnable reference architecture for durable API ingestion and event processing.
- Kafka-compatible streaming with documented partition-key strategy
- Idempotent producers and consumers, bounded retries, and dead-letter handling
- PostgreSQL projections, Redis coordination, and OpenTelemetry instrumentation
- Docker Compose, Kubernetes manifests, CI, and a reproducible k6 load-test harness
API clients -> FastAPI ingestion -> Event stream -> Partitioned workers
| |
Redis idempotency PostgreSQL + DLQ + telemetry
Repository | Architecture decisions
Deterministic CLI for auditing AI and LLM repositories before deployment.
- Checks evaluation, observability, guardrails, security, reliability, RAG quality, and cost controls
- Produces JSON, Markdown, and SARIF output for local use and CI workflows
- Converts production-readiness requirements into reviewable repository evidence
Repository | AI audit practice
Multi-agent engineering workflow with specialized roles and explicit verification stages.
- Separate planning, backend, database, security, testing, and review responsibilities
- Structured handoffs instead of an unbounded prompt loop
- Reusable FastAPI architecture and delivery skills
I submitted open-telemetry/opentelemetry-python-contrib#5111 to ensure Psycopg2 tracing remains active when applications pass a custom cursor_factory per cursor.
The contribution includes focused regression coverage, package-level validation, a changelog entry, and Linux Foundation CLA completion. It is currently open and awaiting maintainer review.
I list upstream work by its real status: submitted while under review, merged only after maintainers merge it.
| Product | System | Engineering focus |
|---|---|---|
| Wellows | AI search visibility platform | Multi-agent workflows, retrieval, citation scoring, and production observability |
| Savyour | Fintech and merchant platform | Distributed services, ledger workflows, partner APIs, and settlement processing |
| EFU Life | Regulated insurance workflows | Auditable workflow services, enterprise integration, and access controls |
| ClassFlow | Live marketplace SaaS | Concurrency control, Redis coordination, matching, and payment workflows |
Detailed outcomes and project context are available in the case studies.
| Area | Tools and practices |
|---|---|
| AI systems | LangGraph, LangChain, OpenAI, Anthropic, Gemini, Qdrant, hybrid search, evaluation, guardrails |
| Backend | Python, FastAPI, Java, Spring Boot, TypeScript, PostgreSQL, Redis |
| Distributed systems | Kafka-compatible streams, SQS, RabbitMQ, idempotency, retries, DLQs, partitioning |
| Operations | AWS, Docker, Kubernetes, Terraform, GitHub Actions, OpenTelemetry, CloudWatch |
| Architecture | ADRs, API contracts, threat modeling, load testing, SLOs, failure-mode analysis |
- Make system boundaries and failure modes explicit.
- Measure retrieval and model behavior instead of relying on demos.
- Keep nondeterministic AI behind deterministic controls.
- Design retries, idempotency, and observability before incidents require them.
- Publish claims only when the repository, test, or case study can support them.
- Improving the production RAG and event-platform reference architectures
- Contributing a Psycopg2 instrumentation fix to OpenTelemetry Python Contrib
- Researching a second upstream contribution without duplicating active work
- Writing about production AI architecture, evaluation, and reliability
I work with teams moving AI prototypes into production and with engineering organizations that need stronger architecture, reliability, or delivery controls.
Portfolio | Case studies | LinkedIn | Book a call | Email
Production AI | Grounded RAG | Distributed Systems | Reliable Delivery