About
I enjoy moving between concrete problems and the abstractions that connect them.
I'm an AI engineer with a mathematics background. I build LLM systems evaluation-first: deterministic where a model isn't needed, measured against ground truth where one is, and open about the error that remains. My flagship project, SteamLens, is a live LLM app that analyzes Steam reviews and cites its own measured accuracy, evaluated against a human-adjudicated gold set before launch. I'm now building an agent system evaluated the same way, inside a synthetic organization whose true answers are known by construction.
Based in Turkey (UTC+3) and currently looking: open to remote roles (contract or EOR) and to relocation. Email is the fastest way to reach me; the resume is a two-page PDF.
How I work
What interests me most isn't just finding an answer, but understanding the structure underneath it: what assumptions it depends on, what information it preserves, and whether the conclusions are actually justified by the evidence. In practice:
- A number with no baseline and no error bar isn't a result yet.
- A conclusion is earned by the experiment that could have falsified it.
- Negative results are findings, and I publish them as such.
- Published runs carry their seed, config and code version, and the model outputs they used are stored alongside them.
- Complexity must earn its place: no framework or layer that doesn't demonstrably buy something.
- Claude Code daily as a design and drafting partner: I work through system shape with it, make the decisions, then build incrementally; each change is accepted only after tests, CI and expected behaviour pass.
How I got here
I studied mathematics at Boğaziçi University (a B.Sc. followed by graduate coursework) and worked as a teaching assistant, leading problem-solving sessions across seven undergraduate courses. Mathematics set the standard I still hold work to: a claim isn't finished until the reasoning behind it holds.
Then came three years as a game developer, building and shipping 20+ production games and interactive experiences for platforms like AirConsole, CrazyGames, and Snapchat. Games taught the other half of the job: real users, real deadlines, and systems judged in production rather than on paper.
AI engineering combines the two: production software whose behaviour is an empirical question. The field calls that discipline evaluation; it's the part I care about most.
Current areas of focus
- LLM and agent systems, and their evaluation
- Machine learning and reinforcement learning
- Data analysis and statistical reasoning
- Experiment design and simulation
- Production AI systems: deployment and operation
I build in public: projects and technical reports here, and posts on LinkedIn and X, including the mistakes, dead ends, and lessons learned.