I am the founder of tiyuvta, an independent AI lab, and a systems software engineer at AWS ElastiCache. My route into research runs through implementation: reproduce the baseline, expose the hidden assumption, train the missing comparison, and let the result change the plan.
tiyuvta is where the research becomes a product, with one job: making inference faster, tokens cheaper, and intelligence better. It has three legs. A hosted inference API for agent workloads at inference.tiyuvta.ai, OpenAI- and Anthropic-compatible, streaming, with speech-to-text on the same key. Setting up inference systems for teams that run their own, locally or in their cloud. And the research itself, measured on the same serving path customers use.
I study computer science at The Open University of Israel while maintaining open-source systems and running research outside a traditional lab. That path has made method unusually important to me. Claims need matched baselines, saved artifacts, explicit noise floors, and a visible record of negative results.
The research is trained, not only measured: compact MTP draft heads with their own vocabularies, healed pruned MoEs, and a fine-tuned ColBERTv2 retriever scored by a trained LoRA relevance judge that beat its prompted 27B teacher. Before focusing on efficient inference I worked deeply in datastores, language clients, queues, retrieval, and developer tools. I still maintain Valkey GLIDE, stay active across the valkey-io ecosystem, and help build desktop developer tools; that record (review, releases, support) runs alongside the studies. Model research becomes real software surprisingly quickly.
I also volunteer as a mentor with Yotzim LaShinui, helping a junior developer from a formerly ultra-Orthodox background find their way into software.