I build reliable AI agents, agent-native infrastructure, and engineering systems that teams can actually run.
I am focused on the operating layer around AI agents: execution reliability, audit trails, tooling distribution, observability, and AI-native development workflows.
- Building durable execution systems for production AI agents.
- Designing AgentOps infrastructure: traces, metrics, logs, replay, governance, and human approval.
- Turning team workflows into skills, prompts, pipelines, and reusable agent systems.
- Shipping practical developer tools that make agents easier to inspect, trust, and operate.
- Bringing 10+ years of backend architecture and reliability experience into AI-native products.
| Area | What I Care About |
|---|---|
| AI agents | Durable execution, task decomposition, tool use, memory, replay, and approval loops |
| AgentOps | Observability, auditability, cost visibility, failure recovery, and production safety |
| Developer tools | CLI-first workflows, skill distribution, local-first automation, and AI coding pipelines |
| Reliability | Graceful deploys, release gates, monitoring, alerting, backup, and incident response |
| Backend systems | Java / Go / Python, distributed systems, high concurrency, cloud-native architecture |
Before moving deeply into AI-native systems, I spent years building and operating backend platforms where reliability was not optional: payments, booking, marketing campaigns, observability platforms, business middle platforms, and high-concurrency services.
That history shapes how I build AI systems now:
- I treat agents as production systems, not just prompts.
- I care about audit trails, rollback paths, monitoring, and data safety.
- I prefer workflows that teams can repeat, review, and improve.
- I still enjoy hard backend problems: concurrency, consistency, storage, release safety, and operational visibility.
Durable execution and reliability runtime for production AI agents.
- Tracks agent execution as auditable ledgers.
- Makes long-running agent workflows easier to replay, inspect, and recover.
- Focuses on reliability, governance, and operational visibility instead of demo-only agent behavior.
A distribution layer for prompts, agent scripts, and MCP servers.
- Treats agent capabilities as versioned, installable assets.
- Supports local runtime setup, metadata indexing, and repeatable installation.
- Designed for teams that need shared AI capabilities without copy-paste drift.
A CLI that gives agents web reach across public sources such as GitHub, Reddit, YouTube, Bilibili, XiaoHongShu, and more.
- Built for agents that need external context without bespoke API integrations.
- Optimized for practical search and reading workflows.
Interactive knowledge graphs for codebases.
- Turns repositories into explorable graphs.
- Helps developers and agents understand unfamiliar systems faster.
- Built around the idea that graphs should teach, not just impress.
An AI coding agent environment for the terminal.
- Focuses on local harnesses, tools, LSP, subagents, and safer edit workflows.
- Explores what a practical AI-native developer environment should feel like.
- Backend architecture: 10+ years designing distributed systems, business platforms, and high-availability services.
- Reliability engineering: observability, alerting, release gates, graceful startup/shutdown, backup and recovery.
- Team leadership: 7+ years across technical governance, engineering workflow design, and cross-functional delivery.
- Production scale: hands-on experience with high-concurrency systems, large traffic events, and multi-team platforms.
- AI engineering: agents, RAG, MCP, LangGraph-style workflows, tool orchestration, and AI-assisted development pipelines.
- Orange Digital Technology: led architecture work and built an observability stack with Prometheus, SLS, Grafana, and N9E; pushed monitoring coverage, alerting, and reliability governance.
- Huanxin Network: built a business middle platform from scratch, supporting multiple business lines and large-scale daily traffic.
- JD Technology: worked on large-scale marketing and coupon systems for major promotional campaigns.
- VIPKID: led booking and class-management architecture, coordinating with multiple teams around high-traffic scheduling scenarios.
- Duolabao: designed core payment systems, OAuth2 open platform capabilities, data synchronization middleware, and data platform work.
- Scalable architecture: distributed systems, high-concurrency services, and business platform design.
- Reliability first: observability, incident response, graceful deploys, release gates, and backup/recovery.
- AI infrastructure: agents, RAG, MCP, AgentOps, skill distribution, and AI coding pipelines.
- Product-minded engineering: turning fuzzy product requirements into systems that can ship and be operated.
- Team enablement: making engineering workflows explicit through docs, checklists, automation, and reusable tools.
- GitHub: github.com/yaogdu
- Email: yaogdu@gmail.com
- Open to discussing: AI agents, AgentOps, developer tooling, backend reliability, and AI-native engineering workflows.