Your employees already use AI.
Today that means private accounts, company data in US clouds, and a bill nobody can explain. ALPCRUN.CH puts one gateway in front of all of it: the models they ask for, served in Europe, with a cost line for every team. Set up in days, not quarters.
Shadow AI is already in your building.
Someone pasted the quarterly numbers into a private ChatGPT account last week. Not out of malice, out of convenience. A policy PDF will not fix that. Something better, easier, and under your control will.
Live in an afternoon
No AI team, no identity project, no six-month pilot. Choose your models, set a budget, invite your teams. Most companies have their first department working before the end of the day.
Every euro accounted for
See exactly what Sales, Legal, and Engineering spend on AI, and cap it before it surprises you. One invoice instead of forty personal subscriptions on expense reports.
Your data stays in Europe
European providers by default, or GPUs reserved for your company alone when the workload demands it. Nothing is used to train anyone's model, so the question from your data protection officer has a one-sentence answer.
Everything IT needs to say yes.
The unglamorous parts of an AI rollout, handled for you: access, permissions, budgets, logs, and a bill that makes sense. All of it shipping today.
Works with the tools they already use
Chat apps, coding agents, your own software. Anything that speaks the OpenAI API points at one URL and is governed from the first request.
No identity project
Your admin issues a dedicated key per person or team, in seconds. No directory sync, no OIDC, no seat licenses to true up.
Teams and departments
Group people the way your company is actually organized, then budget and report along the same lines.
Cost per team, live
Consumption attributed to a team, a project, or a person before it ever reaches your invoice.
Budgets and caps
Ceilings per company, team, or key, by day, week, month, or lifetime. Nobody can spend a surprise into existence.
Access you can revoke
Every key is scoped to the models it may use, with rate limits and an expiry date. Cut it off the day it should stop.
Coding agents, governed
Your developers burn the most tokens by far. Per-person keys, rules that catch a secret before it leaves, and a ceiling that stops a runaway loop.
Change models without a migration
Your people ask for “the everyday assistant.” You decide what actually answers, and swap it later without anyone changing a thing.
A complete audit trail
One row per request: who asked, which model answered, what it cost. Message contents only if you switch that on.
Hosted in the EU
European providers by default, your own dedicated GPUs when you need them. Your data stays in EU regions, and nothing is used to train anyone's model.
Guardrails on the way out
Catch a personal name, a German ID number, an IBAN, or a project codename before it reaches a model. Block it, redact it, or just log it and tell you.
Want to see it on your own company?
Register for early access →Start fast. Scale private.
Most companies should not buy GPUs to find out whether AI is useful. Start on shared European endpoints and stay there as long as they serve you. Move to hardware of your own when a regulator or a workload makes it necessary. Both run in EU regions, and the platform above them does not change.
European token APIs
Nothing to provision. Your teams work the same afternoon and you pay for what they actually used.
Your own private cluster
Dedicated GPUs in the region you choose, for work a regulator will scrutinize or a shared endpoint cannot keep up with.
Monday morning to company-wide.
Three steps, no consultants, no data science hires. You need one person who can click through a dashboard and one person who can decide a budget. They can be the same person.
Set it up
Choose where your data lives, which models people get, and what it is allowed to cost.
Hand out access
Teams and keys in the shape your company already has. People point whatever they use at one endpoint.
See it working
Adoption and spend, side by side, so the AI conversation in your next board meeting has numbers in it.
An answer to “what is AI costing us?”
Almost nobody can answer that question today, because the spend is scattered across personal subscriptions and expense reports. Here it is one invoice, split the way your cost centers are.
# one invoice, split by cost center
Engineering 68 people 12,940 requests €438.20
Legal 11 people 1,180 requests €124.10
Sales 42 people 2,260 requests €96.80
Marketing 19 people 1,940 requests €71.40
HR 9 people 410 requests €29.30
Management 6 people 90 requests €9.10
total €768.90
per employee €3.59
- ▸Chargeback finance can book: consumption rolled up by team, project, or person, exported to the right cost center.
- ▸Caps before surprises: a monthly ceiling per team means the invoice never arrives as news.
- ▸One invoice, not forty: replaces the personal ChatGPT and Copilot subscriptions currently hiding in expense reports.
- ▸Proof of value: usage next to cost, so you can see which departments are getting something out of it.
- ▸Your keys or ours: keep your existing vendor contracts and let us govern and meter them, or take every model on a single invoice from us.
- ▸Priced the way it runs: per token on a shared endpoint, per GPU-hour and traffic on your own cluster. No per-seat guessing either way.
- ▸Euros, exactly: costs are booked in each provider's own currency and rolled up at daily ECB reference rates, fixed per day, so last quarter's totals never move.
Your data never leaves Europe.
Not a checkbox somebody has to remember to tick. It is simply where the hardware is. If you ever want a US frontier model, you switch it on deliberately, for one team and one job.
- ▸Served inside the EU: European providers, or GPUs reserved for your company alone. Either way the request does not leave.
- ▸Never used for training: your prompts and documents answer your questions and do nothing else.
- ▸GDPR without the essay: one region, one processor, and a data processing agreement you can actually sign.
- ▸EU AI Act and DSGVO homework: one-click PDF audit reports mapped to the articles your auditor will cite, built from the record of which model answered what.
- ▸Frontier models, still in Europe: Claude and GPT-class models can be served from EU regions rather than a US endpoint. Sending a request outside Europe stays an explicit, per-team decision.
- ▸German and English: the admin console and the end-user documentation come in both, and your spend reads in euros.
Models are a menu, not a strategy.
Nobody should have to follow model releases to run a company. Your people pick the job they need done, we keep the menu current, and you decide which options are on it.
- ▸Open models, either tier: Mistral, Llama, Qwen, Kimi, Gemma, DeepSeek, and friends. A small model on a single GPU, or one of the biggest that needs multiple racks of them. Sizing the hardware is our job, not yours.
- ▸Your own fine-tunes: upload your weights into the private registry we host and serve them next to the rest.
- ▸Frontier models on request: GPT and Claude attach per model and per team, always labeled as leaving Europe.
- ▸No homework for your staff: a sensible default handles everyday work, so nobody compares benchmarks to write an email.
The details, for the people who ask for them.
Under the friendly surface is a real platform: our own gateway in front of whatever backend suits the workload, European provider endpoints, dedicated GPU clusters running hosted llm-d, or a frontier API, all governed, metered, and priced in one place. Your developers get the same platform your colleagues get.
Private dedicated clusters
- Per-customer GPU clusters, never shared queues
- Created and torn down in minutes
- One GPU up to multi-node for the largest open models
- Scale to zero when idle
- Your models, your KV-cache, your GPUs
Smart request routing
- KV-cache-aware and load-aware steering
- Disaggregated prefill and decode pools
- Multi-region clusters for failover
- More throughput from the same hardware
Your own control plane
- One instance per company, never a shared control plane
- Organizations, teams, and virtual keys beneath them
- Per-key model allow-lists, RPM and TPM limits, expiry
- Priority-tier routing, circuit breaker, bounded failover
- Your vendor keys or ours, sealed per endpoint
Rules before the model sees it
- PII detection for names, German IDs, IBAN, card, email, and phone
- Regex and keyword deny-lists, model allow and deny
- Block, redact, or warn, bound to org, team, or key
- Token and cost ceilings that stop a runaway agent
One row per request, always
- Identity, routing decision, tokens, and cost
- Guardrail verdicts and per-stage latency
- One-click PDF reports mapped to EU AI Act and DSGVO
- Message bodies only when you enable them
- Written asynchronously, never on the request path
OpenAI-compatible API
- Change one base URL, keep your code
- Chat, embeddings, images, audio, and moderations
- Chat frontends and coding agents, unchanged
- Streaming and tool calls included
- Works with existing SDKs and frameworks
Metered on what you run
- Per token on shared endpoints, per GPU-hour on your own
- Itemized per cluster, team, and key
- Spend caps per team or environment
- Low flat baseline when scaled to zero
Private model registry
- Upload your own weights from the dashboard
- Versioned and encrypted, hosted by us
- Pulled only by your own cluster
- Runs beside public models in one pool
Leaving costs you a base URL
- The OpenAI API you already write against
- Open-weight models, on llm-d and vLLM, that run anywhere
- Guardrail rules export as YAML or JSON
- Your fine-tuned weights stay yours
# OpenAI-compatible. Any model. Your own GPUs, in Europe.
$ curl https://inference.alpcrun.ch/v1/chat/completions \
-H "Authorization: Bearer $ALPCRUN_KEY" \
-d '{ "model": "mistralai/Mistral-Small-3.2-24B",
"messages": [{ "role": "user", "content": "…" }] }'
< x-alpcrun-region: eu-central (Germany)
< x-alpcrun-cluster: c-9f2a (private, dedicated)
< x-alpcrun-team: engineering (charged back)
< x-alpcrun-billing: gpu-time (no token charge)
Bring AI to your whole company.
Register for early access and we set it up with you: models chosen, budgets set, teams invited, data in Europe. Usually an afternoon, and you will know what it costs before you commit.
Going to Bits & Pretzels 2026 in Munich? Come by the booth for the two-minute demo.