Edgee On-Premise: The Gateway, Behind Your Firewall
The same compress, route, and observe gateway, now deployable entirely inside your own infrastructure.
Welcome to the official blog of Edgee where our founders and leading thinkers dive deep into the transformative world of edge computing. Here, we explore the latest trends, share groundbreaking innovations, and offer our perspective on the evolving digital landscape. From technical deep dives and industry analyses to visionary outlooks, Edgee Blog is your go-to source for thought leadership in edge computing. Join us as we chart the course towards a more connected, efficient, and innovative future.
The same compress, route, and observe gateway, now deployable entirely inside your own infrastructure.
Compressor V2 is the combination of three independent compression strategies: brevity, tool surface reduction and tool result trimming. Each targets a different layer of an agent's request. This post measures end-to-end the gains of this new composed strategy, on real coding and tool-use workloads with paired statistical tests. The results show a combined 50% per-task cost reduction.
We benchmarked Codex alone against Codex routed through Edgee's compression gateway on the same repo, with the same model, under the same workflow. The result: Codex + Edgee used 49.5% fewer input tokens, improved cache hit rate from 76.1% to 85.4%, and reduced total session cost by 35.6%. This post breaks down why context compression makes Codex more efficient, more frugal, and materially cheaper to run without sacrificing useful output.
Today we're shipping two things: Codex support for the Edgee compressor, and Session Reports.
LLM providers go down, hit rate limits, and time out. Here's how Edgee handles retries, fallbacks, and provider scoring to keep your requests succeeding, transparently.
We ran a head-to-head endurance test: raw Claude Code vs Claude Code with Edgee's token compressor. Same plan, same tasks. One went 26.5% further.
Bring Your Own Keys (BYOK) lets teams use their own provider API keys with Edgee while benefiting from token compression, routing, observability, and usage tracking.
A short tutorial video walking you through the essentials of Edgee AI Gateway — what it is, how to get started, and how to route and manage your LLM traffic in under two minutes.
Edgee AI Gateway introduces a new operational layer for production LLM systems, built to reduce costs while bringing visibility and control to AI infrastructure.
AI costs are rising. The root cause isn't economic—it's operational. Cost observability, meaning real, request-level, attributable cost tracking, is the missing layer. And without it, every other optimization strategy is flying blind.
Would you like to find out more about Edgee, test our services or our upcoming features? We’d love to hear from you. Please fill in the form below and we’ll be in touch.