Product at the intersection of AI and business. I define what gets built, then build it: scope, architecture, and execution pipeline.
Based in San Francisco. Open to product roles at AI companies.
Below, two kinds of thing. Products are opened and used by a person. Integrations sit in the path of other software and earn their place invisibly.
AI screen assistant · github.com/M19K/handrail · Product · Apache-2.0 · Windows and macOS
A desktop overlay that sees your screen and walks you through complex software, one step at a time. You ask inside the app you are already stuck in. It looks at what is actually in front of you and points at the real button, not the one in a tutorial.
A real question, a real screen, a real arrow. Captured from the shipped build, not a mockup.
Built for people who can follow an instruction precisely but cannot diagnose why step 4 failed.
- Sees your screen. Every question carries a screenshot, so answers are about your situation rather than the software in general.
- Points at things. An on-screen arrow lands on the actual control.
- Keeps up. For multi-step tasks it watches for each step and moves on by itself.
- Tells you when you have gone wrong. The reason it exists.
Built with: Electron, a vision LLM over OpenRouter, and a step-watching loop that re-reads the screen between turns.
Design decision I'd defend: the API key is encrypted in the OS keychain and never crosses into the UI layer after setup. Only a masked hint does. Screenshots and conversations never leave the machine, and there is no telemetry, no account, no sign-up. A screen assistant that phones home is a different product, and a worse one.
AI career clarity companion · caseyai.co · app.caseyai.co · Product · Founder
An emotionally intelligent, voice-first career coach for young adults navigating career uncertainty. Users talk through their background and blockers across several sessions. CASEY builds understanding conversationally, then converts it into a committed direction and an actionable roadmap.
- Session 1, Discovery. Understand situation, interests, blockers.
- Session 2, Orientation. Identify direction, assign persona.
- Session 3, Roadmap. Generate a three-step plan: Explore, Validate, Confirm.
- Session 4 onward, Accountability. Progress check-ins.
Built with: Hume AI EVI for empathic voice (speech to text, LLM, expressive text to speech), Next.js 15, Node and Express, Supabase with row-level security, OpenAI.
Design decision I'd defend: session prompts and phase configuration live in database tables rather than application code. Coaching behaviour is tunable without a redeploy, and the admin surface exposes it directly, which matters when the prompt is the product.
Shared memory for several AI agents · github.com/M19K/mikoshi · Integration · MIT
Agents forget. Every session starts blank, so the same decision gets re-made and the person in the middle ends up being the memory. Mikoshi makes a folder of plain markdown behave like shared memory.
- A protocol. Who may write what, and how it gets recorded.
- A channel. Agents ask each other for things instead of going through you.
- An ingestion funnel. RSS, YouTube, X, Instagram, newsletters, mail, meeting transcripts and Drive, on a schedule.
- Answers you can check. Every sentence is pinned to a file and a line you can open. Sentences it cannot pin are deleted before you see them.
Measured on a fixed 30-question set: cites the correct file 85% of the time, 87% of sentences supported by the line beside them, zero invented citations across every run.
Built with: Python, SQLite with sqlite-vec, Ollama for local inference, MCP. No database server, no cloud, no API key required.
Design decision I'd defend: it runs on a local model by default and a paid one only when you opt in, per run. Switching to a paid route is a money decision, and money decisions should never be a default.
Measure what a cheaper model actually costs you · github.com/M19K/superrouter · Integration · Apache-2.0
Every routing tool picks a model on price and latency, which are the two things that cannot tell a fluent wrong answer from a right one. SuperRouter measures quality per sub-task on golden sets generated from your own material, then routes on the measurement.
- The finding it exists because of. "QA" is not one job. One model caught 83% of planted screen defects and hit 11% of click targets: perfect at describing a screen, effectively blind at pointing.
- What that costs you. Route on a task label that coarse and you build an agent that describes beautifully and clicks at random.
Built with: Python, OpenRouter, and an OpenAI-compatible local proxy so any agent can point at it without an adapter.
Design decision I'd defend: it ships as a method aimed at your product, never as a published routing table. Measured across two products, which model is better mostly transfers. How good any of them is does not, dropping a median 22 points on work it was not measured on. A leaderboard would be selling the half that does not travel.
Open-source AI teammates · Product · Apache-2.0 · pre-build, repo opening soon
Free, open-source AI teammates you can hand real work to.
- A persistent agent computer. Browser, filesystem, terminal, shared by all your agents.
- Bring your own API key. No account, no central service.
Design decision I'd defend: no hosted tier and no account. The agent computer runs as a container you own, and the model plane is provider-agnostic, so the thing you are trusting with a filesystem and a shell is a thing you can inspect and turn off.
Courses into a queryable knowledge base · github.com/M19K/course-distiller · Product · MIT · local-first
Turn a course you have access to into a private, LLM-queryable knowledge base, on your own machine, for $0 in API spend.
- Drives your own logged-in browser session, so it reaches what you are entitled to reach.
- Transcribes video with Whisper and distills each transcript through a local model via Ollama.
- Pulls text out of attachments: PDF with OCR fallback, xlsx, pptx, docx.
- Compiles the lot into per-lesson markdown and NotebookLM-ready bundles.
Two things I'd point at. Nothing leaves the machine: no API key, no course content, no telemetry, because the expensive parts run locally. And the .gitignore is written so extracted material can never be committed. The tool is for content you are entitled to access, not for redistributing it.
Built with: Python, Playwright, Whisper (MLX on Apple Silicon), Ollama, Streamlit.
TypeScript · JavaScript · Next.js · React · Electron · Node · Python · PostgreSQL · Supabase · Tailwind
AI: vision models · Hume EVI · OpenAI · Anthropic · OpenRouter · Ollama and local inference · retrieval-augmented generation · MCP and agent frameworks · model evaluation and routing
How I work: measure before setting a threshold; record what was turned down and why, not only what was chosen; quote a range rather than a run, because a single measurement has no noise floor.
maazkazi.com · LinkedIn · X · caseyai.co