# Lithus: We run your infrastructure so you don't have to

Managed Kubernetes on bare metal in Germany. ~50% cheaper than AWS. Predictable monthly costs, dedicated hardware, and an SRE team that gets paged at 3 AM so you don't.

## Embedded SRE Team

2 SRE days per €5,000/month. No separate consulting fee. SRE time is built into the same price as your hardware and management. Not ticket-queue hours. Staff-level engineers who know your stack.

- 2-hour incident response SLA
- On-call 24/7: we get paged, not you
- Slack Connect for day-to-day communication

## We Carry the Pager

If Kubernetes, PostgreSQL, Kafka, or any component in our stack breaks at 3 AM, our on-call engineer fixes it. We own the platform, you own the application.

- Kubernetes and third-party software upgrades
- Database failover and backup verification
- Incident response within 2 hours

## Leave It With Us

We handle the full migration: cluster design, hardware provisioning, workload migration, and cutover. 4-6 weeks from order to production-ready cluster, then phased migration at your pace. Your existing infrastructure stays live until you are confident.

- No billing until your workloads are running
- We run the migration, not your team

## What you get

A production-grade Kubernetes cluster on dedicated bare-metal servers, primarily hosted in Hetzner's German data centres. Cilium networking, Mayastor and ZFS storage, a full Grafana/Prometheus/Loki observability stack, and an SRE team embedded in your workflow.

We have been building and operating this stack since Kubernetes 1.0, for SaaS platforms, IoT telemetry pipelines, and data processing workloads.

Engagements start at €5,000/month. Dedicated hardware, full infrastructure stack, and an embedded SRE team. Billing starts when your workloads are running.


---

# Services & Stack

A standard infrastructure stack, managed end-to-end and customised to your workload. We use the same core tooling across every cluster, but the configuration, sizing, and additional services are tailored to what you actually run.

## Infrastructure Stack

Every cluster is built from these core components, configured and customised to your workload. This is the baseline. During planning, we add whatever else your workload requires. Open-source throughout, [no vendor lock-in](/faq#is-there-vendor-lock-in). If you leave, you take the whole setup with you.

### Networking & Ingress

Requests routed through Cilium's eBPF dataplane. Sub-millisecond inter-node latency across a dedicated fibre network, network policies enforced by default.

Tools: Cilium CNI, cert-manager, External DNS, Cilium Network Policies

### Kubernetes on Bare Metal

Workloads scheduled across dedicated bare-metal nodes via RKE2, connected by a dedicated fibre network with NVMe storage on every node. Works with your existing CI/CD and deployment tooling. KEDA handles event-driven autoscaling.

Tools: RKE2, Dedicated Fibre Network, NVMe, Helm, KEDA

### Managed Stateful Services

Databases, caches, and message brokers deployed and managed as part of the cluster. HA configurations, connection pooling, automated snapshots to offsite EU-based S3.

Tools: PostgreSQL (StackGres), MySQL, Valkey/Redis, Kafka, RabbitMQ, Elasticsearch, ClickHouse, MongoDB, + anything your workload needs

### Storage

Replicated block storage, local NVMe, and S3-compatible object storage on dedicated drives. Configured per workload: large slow storage or fast NVMe, encrypted at rest.

Tools: OpenEBS Mayastor, ZFS LocalPV, MinIO/Garage/Ceph, JuiceFS

### Metrics, Logs, Traces

Every layer instrumented. Metrics collected by Prometheus, logs aggregated by Loki, traces via OpenTelemetry and Tempo. Alerts fire to Slack or PagerDuty with a 2-hour incident response SLA.

Tools: Prometheus, Grafana, Loki, Tempo, OpenTelemetry, Alertmanager, Beyla eBPF

### Security & Connectivity

WireGuard encryption between nodes, SSO via Keycloak, intrusion detection via CrowdSec. Mesh VPN and site-to-site tunnels for hybrid connectivity.

Tools: Keycloak, CrowdSec, WireGuard, Headscale, Tailscale

### Build & Ship

Container registry, Git forges, and CI runners, all on your own infrastructure. Firecracker microVMs for fast, isolated build environments.

Tools: Harbor, Forgejo, GitLab, GitHub Actions, Forgejo Actions, GitLab CI, Firecracker

### Backup & Recovery

Cluster-level snapshots, file-level backups, and database point-in-time recovery. Offsite replication to geographically separate EU storage on independent providers. Retention: hourly, daily, weekly, monthly.

Tools: Velero, Restic, PostgreSQL WAL Streaming, Point-in-Time Recovery

## Support & Operational Model

### SRE Allocation

2 dedicated SRE days per €5,000/month. No separate consulting fee. SRE time is built into the same [flat monthly price](/pricing) as your hardware and management.

### Incident Response

2-hour response SLA for infrastructure incidents. We carry the pager for the full stack. If Kubernetes, PostgreSQL, Kafka, networking, or storage breaks, we fix it, day or night.

### Communication

Direct chat (Slack, Teams, whatever you use) for day-to-day communication. Optional shared task board for tracking work. Monthly infrastructure review call. No ticket portals, no chatbots.

## How We Run Migrations

We've migrated workloads off AWS, GCP, and other managed providers. The specifics change, but we follow the same general structure every time. See real examples in our case studies: [PrepBusiness](/customer-success/prepbusiness) (45-day migration) and [Futurepump](/customer-success/futurepump-iot) (Google Cloud to bare metal).

1. **Cluster Design**: We audit your current infrastructure, map your workloads, and design a target cluster topology. You approve the design before we order hardware.
2. **Provisioning (4-6 weeks)**: Hardware procurement, burn-in, and rack installation at the data centre. Kubernetes cluster bootstrap, networking, storage, and observability stack deployment. Full CI/CD pipeline setup.
3. **Migration Execution**: Dry runs until the process is nailed down, then the real cutover. We have automated tooling for migrating out of RDS, Supabase, and other managed providers. Your existing infrastructure stays live until everyone is confident.
4. **De-Risked Billing**: No setup fee. Monthly billing starts when your workloads are running. For larger clusters, we may invoice the first few months in advance to cover hardware procurement.


---

# Pricing: Predictable Costs, Unlimited Egress

Flat monthly rate. Hardware, Kubernetes, observability, and SRE time, all included. Engagements start at €5,000/month. No maximum.

## Pricing Tiers

Three presets optimised for different workload profiles:

### Compute

More CPU cores and memory, less storage. For data processing, simulations, and API-heavy applications that need sustained compute.

### Balance

Even split of compute and storage. Fits most general-purpose workloads: web applications, background processing, databases.

### Storage

More disk capacity, fewer CPU cores. For large datasets, media processing, log aggregation, or backup repositories.

All tiers are ~50% less than equivalent on-demand infrastructure on AWS, Google Cloud, or Azure (excluding egress fees and managed service surcharges). Pricing is illustrative. We size each cluster to your workload profile. Egress is unlimited under applicable fair use policy.

## What's included

- **Dedicated Hardware**: Bare-metal servers, primarily in German data centres. No shared tenancy, no noisy neighbours.
- **Embedded SRE Time**: No separate consulting fee. SRE time is built into the same price as your hardware and management.
- **Full Observability Stack**: Grafana, Prometheus, Loki, and Alertmanager. Dashboards and alerts included from day one.
- **EU Data Sovereignty**: Data processed and stored in Germany by default, on hardware we control. No US hyperscaler, no CLOUD Act ambiguity.

## How to explain this to your CFO

| CFO concern | How we address it | Bottom-line impact |
|---|---|---|
| "Our cloud bill is unpredictable" | Flat monthly rate. No usage metering, no egress fees | One budget line you can forecast 12 months out |
| "We can't find or keep DevOps engineers" | SRE team included. No separate consulting fee | Infrastructure expertise without a full-time hire |
| "Performance issues eat developer time" | Dedicated bare metal, private network. No shared tenancy | Your developers ship product instead of debugging infrastructure |
| "What happens if we want to leave?" | Standard Kubernetes. No proprietary services or APIs | Workloads move to any Kubernetes provider. No exit penalty |
| "How does cost scale as we grow?" | Linear. More hardware, proportionally more spend | Finance can model next year's infrastructure cost in a spreadsheet |
| "What does migration actually cost us?" | We handle the migration. Billing starts when workloads are running | Phased rollout minimises overlap. No big-bang migration risk |


## Infrastructure by Budget

Representative configurations at selected price points. Same figures as the interactive pricing calculator.

### Compute

| Monthly budget | CPU cores | RAM | NVMe | Storage | SRE days/month | Savings vs cloud |
|---|---|---|---|---|---|---|
| €5,000 | 120 | 1,280 GB | 9,600 GB | — | 2 | ~56% |
| €10,000 | 264 | 2,816 GB | 21,120 GB | — | 4 | ~60% |
| €25,000 | 600 | 6,400 GB | 48,000 GB | — | 10 | ~56% |
| €50,000 | 1248 | 13,312 GB | 99,840 GB | — | 20 | ~58% |
| €100,000 | 2520 | 26,880 GB | 201,600 GB | — | 40 | ~58% |

### Balanced

| Monthly budget | CPU cores | RAM | NVMe | Storage | SRE days/month | Savings vs cloud |
|---|---|---|---|---|---|---|
| €5,000 | 72 | 768 GB | 5,760 GB | 211 TB | 2 | ~55% |
| €10,000 | 144 | 1,536 GB | 11,520 GB | 422 TB | 4 | ~55% |
| €25,000 | 312 | 3,328 GB | 24,960 GB | 968 TB | 10 | ~50% |
| €50,000 | 888 | 9,472 GB | 71,040 GB | 1355 TB | 20 | ~55% |
| €100,000 | 2016 | 21,504 GB | 161,280 GB | 2129 TB | 40 | ~57% |

### Storage

| Monthly budget | CPU cores | RAM | NVMe | Storage | SRE days/month | Savings vs cloud |
|---|---|---|---|---|---|---|
| €5,000 | 72 | 768 GB | 5,760 GB | 211 TB | 2 | ~55% |
| €10,000 | 96 | 1,024 GB | 7,680 GB | 528 TB | 4 | ~50% |
| €25,000 | 192 | 2,048 GB | 15,360 GB | 1355 TB | 10 | ~46% |
| €50,000 | 336 | 3,584 GB | 26,880 GB | 3291 TB | 20 | ~50% |
| €100,000 | 600 | 6,400 GB | 48,000 GB | 6969 TB | 40 | ~50% |

---

# Infrastructure Spectrum

Four models. Different trade-offs. One question: how much operational responsibility do you want to carry?

## 1. Public Cloud (AWS / GCP / Azure)

Fully managed services. You pay a premium for convenience and the ability to scale to zero. Best for: variable workloads that spike unpredictably.

**Pros**: Scale-to-zero, global edge network, managed databases
**Cons**: 2-5x hardware cost, vendor lock-in, opaque pricing, egress fees

## 2. Managed Private Cloud (Lithus)

Dedicated bare-metal servers, fully managed Kubernetes stack, and embedded SRE capacity. You get the cost and performance benefits of your own hardware without hiring an infrastructure team.

**Pros**: ~50% cheaper than public cloud, ~2x bare-metal performance, dedicated SRE team, EU data sovereignty
**Cons**: 4-6 week provisioning lead time, €5k/month minimum, not suitable for scale-to-zero workloads

## 3. Rented Bare Metal (Hetzner / OVH direct)

You rent the servers and manage everything yourself. Lowest per-unit cost, but you carry the full operational burden: OS patching, Kubernetes upgrades, networking, monitoring, on-call.

**Pros**: Cheapest per-unit cost, full control, no abstraction layers
**Cons**: You are the on-call engineer, no SRE support, DIY Kubernetes, hiring required

## 4. Owned Hardware (Colocation)

You buy the servers, rack them in a data centre, and manage everything from BIOS to application. Maximum control, maximum responsibility.

**Pros**: Lowest long-term cost at scale, complete hardware control
**Cons**: Capital expenditure, hardware refresh cycles, physical access required, full ops team needed

## Why Managed Private Cloud Wins for Steady-State Workloads

### Network Latency

Inter-node latency is ~0.2ms. Public cloud instance-to-instance is typically 1-3ms with jitter. For database replication, distributed caches, and service meshes, that difference compounds across every request.

### Dedicated Resource Access

No noisy neighbours. Your NVMe drives, CPU cores, and network interfaces are not shared with other tenants. Consistent IOPS and throughput, no performance variance.

### Architectural Simplicity

One Kubernetes cluster. Standard APIs. No proprietary service mesh, no cloud-specific IAM, no vendor SDK required. Your Helm charts and Terraform modules work here the same as anywhere. If you leave, you take your workloads with you.


---

# Lithus vs. Public Cloud

Side-by-side comparison on the metrics that matter.

If you are evaluating whether to move off AWS, GCP, or Azure, here is an honest comparison. We include the areas where public cloud is the better choice, because knowing when something is not the right fit saves everyone time.

## At a glance

| | Public Cloud (AWS/GCP/Azure) | Lithus |
|---|---|---|
| **Billing Model** | Per-resource, per-hour + egress + managed service fees | Flat monthly rate, all-inclusive |
| **Performance** | Shared tenancy, variable IOPS | Dedicated bare metal, consistent throughput |
| **Support** | Ticket queue, tiered response times | Embedded SRE, direct chat, 2-hour response |
| **DevOps Labour** | You hire or outsource separately | Included: 2 days per €5k/month |
| **Billing Start** | Immediately on provisioning | When workloads are running |
| **Relative Cost** | Baseline (1x) | ~0.5x for equivalent hardware |
| **Scale-to-Zero** | Native support for serverless and scale-to-zero | Not available. Dedicated hardware runs 24/7 |
| **Global Presence** | 60+ regions worldwide, edge network | Primarily Germany (Hetzner). Other providers available on request |
| **Managed Service Breadth** | 100+ managed services (ML, analytics, IoT, etc.) | Standard stack: Kubernetes, PostgreSQL, Kafka, object storage, observability |

## What the table doesn't tell you

### Billing: flat rate, no surprises

The biggest problem with cloud billing is not the total. It is the surprises. One of our clients was paying $20,000-30,000/month on AWS, which was manageable. Then a small configuration change triggered $2,500/day in EFS transfer costs with no warning. As they put it: "Small configuration changes on AWS lead to them charging you orders of magnitude more money."

With a flat monthly rate, your CFO can forecast infrastructure cost 12 months out in a spreadsheet. No usage metering, no egress fees, no managed service surcharges. Use our [pricing calculator](/pricing) to see what your budget buys.

### Performance: dedicated hardware, no noisy neighbours

Cloud instances share physical hardware with other tenants. Your IOPS and throughput vary depending on what your neighbours are doing. On dedicated bare metal, inter-node latency is ~0.2ms across a private fibre network. Public cloud instance-to-instance is typically 1-3ms with jitter. For database replication, distributed caches, and service-to-service calls, that difference compounds across every request in every transaction.

Your NVMe drives, CPU cores, and network interfaces are not shared. Consistent IOPS at 3 PM and 3 AM, regardless of what anyone else is running.

### Support: engineers on Slack, not tickets in a queue

Cloud provider support means a ticket portal with tiered response times. Even on Business or Enterprise plans, you talk to a support agent who manages hundreds of accounts. With Lithus, you get a shared Slack channel with the same engineers who built your cluster. No handoffs, no runbooks, no "let me escalate this."

One client was paying $2,000/month for Supabase when a database lock lasted 72 hours and support could not explain what happened. After moving that workload to us, their CTO said: "Thank you for being so responsive. It is awesome." The difference is not the SLA number. It is whether someone who understands your setup is actually looking at the problem.

### DevOps labour: included, not outsourced

On public cloud, you still need someone to manage Kubernetes upgrades, configure alerting, debug networking, optimise database performance, and handle incidents. That is a full-time hire or an expensive consulting engagement.

With Lithus, SRE capacity is built into the price: 2 engineering days per €5,000/month. These are staff-level infrastructure engineers, not junior support agents. For one client, that included building a custom Kubernetes operator in Rust, deploying a private mesh network, and setting up compliance monitoring. See the [full stack](/services).

### Migration risk: we carry it, not you

On AWS, billing starts the moment you provision a resource. With Lithus, billing starts when your workloads are running on our infrastructure. No setup fee. We handle the migration: cluster design, hardware provisioning, workload migration, and cutover. Your existing infrastructure stays live until you are confident. See our [FAQ](/faq) for typical migration timelines.

### Cost: ~50% less for equivalent hardware

For steady-state workloads, dedicated bare metal costs roughly half of equivalent on-demand cloud infrastructure.

## Where public cloud wins

We are not the right fit for every workload. Here is where AWS, GCP, and Azure are genuinely better.

### Scale-to-zero and spiky workloads

Dedicated hardware runs 24/7. If your workload is highly variable, drops to zero for hours, or you need serverless functions, public cloud is the right answer. Lithus is built for steady-state workloads that run continuously: SaaS platforms, data pipelines, APIs with predictable traffic, databases.

### Global edge presence

Our infrastructure is primarily in Germany, with other providers available on request. If you need 60+ regions, a global CDN, or sub-10ms latency to users in Asia-Pacific, public cloud has infrastructure we cannot match. For most European and transatlantic workloads, a single well-connected location is sufficient.

### Niche managed services

AWS has 200+ managed services. We run a [standard open-source stack](/services): Kubernetes, PostgreSQL, Kafka, Redis, object storage, observability. If your workload depends on a service with no open-source equivalent (SageMaker endpoints, BigQuery, DynamoDB Streams), public cloud is harder to leave. For most workloads, the open-source equivalent works the same or better on dedicated hardware.

Not sure which model fits? Our [infrastructure spectrum](/infrastructure) page compares all four options with honest trade-offs.



---

# About Lithus

Infrastructure engineers who run managed private clouds for SMEs. Operating since 2007. Infrastructure hosted primarily in Germany.

## No-BS Infrastructure

We are infrastructure engineers. We build and operate managed Kubernetes clusters on bare-metal servers in German data centres. Our clients are SMEs and software houses that need production-grade infrastructure without hiring a platform team.

We do not have a sales team. We do not sponsor conferences. We spend our time on uptime, performance, and keeping our clients' infrastructure boring, in the best possible way. If something breaks at 3 AM, we fix it. That is the job.

One client, a climate data analysis company, came to us from public cloud with unpredictable costs and growing spend. We migrated them to managed bare-metal. The result: predictable billing, better performance, and a team they can reach on Slack when something needs attention.

## Operational Ownership

We own uptime for the full stack. Kubernetes, PostgreSQL, Kafka, networking, storage, observability. If we deployed it, we carry the pager for it. When something goes wrong, we get the alert, we diagnose it, and we fix it. You get a postmortem.

We own the platform, you own the application. No finger-pointing, no ambiguous responsibility. If your pods are crashing because of a node issue, that is on us. If they are crashing because of a memory leak in your code, we will tell you and set up temporary mitigations to keep things running in the meantime.

## Direct Access

When you work with Lithus, you work with the same engineers who built the platform. There is no account manager buffer, no offshore handoff, no "let me check with the team."

We are honest about timelines and realistic about scope. Hardware takes 4-6 weeks. Migrations play out over a couple of months depending on complexity. If something shifts, you hear it from us directly with a revised plan, not a vague status update.

## ISO 27001

We are working towards ISO 27001 certification. Our information security management system covers access control, incident response, change management, and supplier oversight. Hetzner's data centres are already ISO 27001 certified and run on 100% renewable energy. For clients with compliance requirements, we provide documentation on request.

## Where We Deploy

Most of our clusters run on Hetzner dedicated servers in Germany — ISO 27001 certified data centres with redundant power, cooling, and networking. The hardware-to-cost ratio is the best in Europe, and we work with their custom solutions team to provide multi-AZ and multi-region clusters.

For our clients, this means: EU-sovereign infrastructure, predictable hardware costs, and no dependence on US hyperscaler pricing decisions. Your data stays in the EU, processed on hardware we control, managed by engineers in our team.

We also deploy on other providers when a workload needs a specific region or configuration. The stack and operational model remain the same.


---

## What is managed private cloud?

Managed private cloud sits between public cloud (AWS/GCP/Azure) and renting your own bare metal. You get dedicated bare-metal servers with a fully managed Kubernetes stack and an SRE team, without hiring an infrastructure team yourself. Lithus handles the hardware, networking, Kubernetes, databases, monitoring, backups, and on-call. You deploy your application and focus on your product.

Four infrastructure models exist on a spectrum (see our [infrastructure spectrum](/infrastructure) page for a full comparison):
1. **Public cloud** (AWS/GCP/Azure): most convenient, most expensive, usage-based billing
2. **Managed private cloud** (Lithus): ~50% cheaper, dedicated hardware, SRE team included, €5k/month minimum
3. **Rented bare metal** (Hetzner/OVH direct): cheapest per unit, but you manage everything yourself
4. **Owned hardware** (colocation): lowest long-term cost at scale, requires capital expenditure and a full ops team

## How much does Lithus cost?

Flat monthly rate starting at €5,000/month. The price includes dedicated bare-metal servers, the full Kubernetes and observability stack, and embedded SRE time. No egress fees, no per-resource metering, no managed service surcharges. Use our [pricing calculator](/pricing) to see what your budget buys.

SRE time scales with spend: 2 engineering days per €5,000/month. So a €10,000/month cluster includes 4 days of SRE capacity. This is not ticket-queue hours; it's staff-level engineers who know your stack.

Typical savings versus equivalent AWS/GCP/Azure infrastructure: ~50%. See our [side-by-side comparison](/compare) for a full breakdown.

## How long does migration take?

4-6 weeks from order to production-ready cluster (hardware procurement, burn-in, Kubernetes bootstrap, full stack deployment). Then phased workload migration at your pace, typically 1-3 months total depending on complexity.

Your existing infrastructure stays live during migration. No setup fee. Billing starts when your workloads are running on our infrastructure, not when provisioning begins. For larger clusters, we may invoice the first few months in advance to cover hardware procurement. For a real-world example, see how we migrated [PrepBusiness with a 45-day turnaround](/customer-success/prepbusiness).

## What's included in Lithus SRE support?

- 2 SRE days per €5,000/month of spend (built into the price, not a separate consulting fee)
- 2-hour incident response SLA for infrastructure issues
- 24/7 on-call: we carry the pager for the full stack
- Direct Slack Connect channel for day-to-day communication
- Monthly infrastructure review calls
- Kubernetes and software upgrades
- Database failover and backup verification
- No ticket portals, no chatbots

To see what this looks like in practice, read how we [diagnosed and eliminated mysterious downtime for PrepBusiness](/customer-success/prepbusiness) or [built custom tooling for XDI Systems](/customer-success/xdi).

## Where is Lithus infrastructure hosted?

Infrastructure is hosted primarily on Hetzner dedicated servers in German data centres. EU data sovereignty by default: no US hyperscaler, no CLOUD Act ambiguity. All data processed and stored on hardware Lithus controls.

Hetzner provides predictable hardware costs and a positive reputation among technical audiences. Lithus adds the managed Kubernetes stack, SRE team, and operational responsibility on top. For clients requiring specific regions or configurations, we also work with other providers. More on our [About page](/about#why-hetzner).

## What Kubernetes distribution does Lithus use?

RKE2 (Rancher Kubernetes Engine 2), deployed on dedicated bare-metal servers connected by a dedicated fibre network with sub-millisecond inter-node latency. The full stack includes:

- **Networking**: Cilium CNI with eBPF dataplane, cert-manager
- **Storage**: OpenEBS Mayastor, ZFS LocalPV, MinIO/Garage/Ceph for S3-compatible object storage
- **Databases**: PostgreSQL (StackGres), MySQL, Valkey/Redis, Kafka, RabbitMQ, Elasticsearch, ClickHouse, MongoDB
- **Observability**: Prometheus, Grafana, Loki, Tempo, OpenTelemetry, Alertmanager
- **Security**: Keycloak, CrowdSec, WireGuard, Headscale/Tailscale
- **CI/CD**: Harbor, Forgejo/GitLab, GitHub/Forgejo/GitLab Actions, Firecracker microVMs
- **Backup**: Velero, Restic, PostgreSQL WAL streaming with point-in-time recovery

Open-source throughout. If you leave, you take the whole setup with you. See our [Services & Stack page](/services) for the full breakdown of each layer.

## What happens if something breaks at 2am?

Our on-call engineer gets paged and fixes it. 2-hour incident response SLA. We carry the pager for Kubernetes, databases, networking, storage, and monitoring. You own your application; we own the platform. See our [full stack breakdown](/services) for exactly what we cover.

## Is there vendor lock-in?

No. The entire stack is open-source and runs on standard Kubernetes APIs. Your Helm charts, Terraform modules, and CI/CD pipelines work here the same as on any other Kubernetes provider. No proprietary services, no vendor SDK, no exit penalty. If you leave, you take your workloads with you. See the [full stack](/services) — it's standard tooling throughout.

## Why not just rent Hetzner servers directly?

You can — and if you have an infrastructure team to manage them, it's a good option. Hetzner gives you cheap dedicated hardware. What it doesn't give you is a production Kubernetes stack, managed databases, observability, backups, security hardening, or anyone to page at 2am when something breaks.

Lithus runs on Hetzner hardware. The value isn't the servers — it's everything on top: RKE2, Cilium, StackGres, Prometheus/Grafana/Loki, Velero backups, and an SRE team that knows your stack. If you have the team to build and maintain all of that yourself, renting direct is cheaper. If you don't, you're comparing Lithus not against Hetzner's price but against Hetzner's price plus 1-2 full-time infrastructure engineers (€80k-€150k/year each).

See our [infrastructure spectrum](/infrastructure) for where managed private cloud sits relative to renting bare metal directly.

## Do I need to hire DevOps engineers?

No. SRE/DevOps capacity is included in the price. 2 engineering days per €5,000/month. These are staff-level infrastructure engineers available via Slack who know your stack, not generic support agents working from a runbook. For most clients in the €5k-25k/month range, this replaces the need for a dedicated infrastructure hire entirely. [Schedule a call](/contact-us) to discuss your specific needs.


---

# XDI Systems: From Unpredictable AWS Bills to Bare Metal

**Industry**: Climate Risk Analytics | **Location**: Australia / UK | **Team**: ~45 people | **Website**: xdi.systems

> "I tell people we moved off AWS, but that doesn't really capture it. We got a team that built us custom tooling, migrated our databases, deployed a private network, and responds on Slack within hours. AWS doesn't offer that at any price." - Tim McEwan, CTO, XDI Systems

## The Situation

XDI Systems provides physical climate risk assessments to major global banks and governments. Their models analyse individual assets (harbours, buildings, critical infrastructure) against climate projections.

Their AWS setup had grown organically over several years: three separate EKS clusters across multiple sub-accounts, legacy servers, a managed database service, and configuration spread across multiple systems. One engineer was managing all of it alongside other responsibilities.

Monthly AWS spend averaged $25,000, with unpredictable spikes. One EFS transfer event cost $20,000 over three days. Legacy on-demand instances for data science work added another $5-10k/month with no cost governance.

## What We Built

Migrated XDI Systems from three separate AWS clusters to a single multi-AZ bare-metal cluster in Germany, running our [full managed Kubernetes stack](/services). Logically separated into production, QA, and development environments. Migration phased over four months: QA first, then production, then development. Billing started when workloads were running.

**Storage**: Dedicated object storage cluster and JuiceFS as a POSIX-compatible shared filesystem, replacing AWS S3, Cloudflare R2, and EFS. The object storage cluster benchmarks at 200 Gbps aggregate throughput handling 50,000 requests per second. At sustained throughput, the equivalent S3 GET request volume alone would cost ~$55,700 USD per month on AWS. On bare metal, the cost is fixed.

**Data scientist VMs**: Moved data science VMs into the cluster using KubeVirt with a custom Kubernetes operator and CLI built in Rust. Full VM lifecycle management: provisioning, networking, resizing, monitoring, and shutdown. Each VM automatically connects to a private Tailscale mesh network via open-source Headscale.

**Database**: Migrated off a managed database vendor into in-cluster PostgreSQL. Every database runs as a primary/replica pair with snapshot backups and point-in-time recovery, with 2-4x the resources of the managed instances they replaced.

**Private network**: Self-hosted Headscale server giving the entire team secure access to the cluster and internal services through a Tailscale mesh network. Simplified compliance and gave everyone access to internal tooling without managing traditional VPN infrastructure.

**Compliance**: Deployed Falco for runtime threat detection at the kernel level and Kyverno for policy enforcement, replacing AWS Security Hub.

## The Numbers

| | Before | After |
|---|---|---|
| Monthly cost | $25k USD, with unpredictable spikes | ~45% less, [fixed monthly rate](/pricing) |
| Clusters to manage | 3 separate EKS clusters | 1 multi-AZ cluster, 3 logical environments |
| Storage | EFS ($2,500/day during spikes) + Cloudflare R2 | 200 Gbps dedicated cluster, flat cost |
| Data scientist VMs | On-demand AWS instances, no cost controls | KubeVirt VMs with CLI provisioning and lifecycle management |
| Database | Managed vendor, resource-constrained. Costs growing with data volume | Primary database self-hosted in-cluster, with 2-4x resources |
| Private network access | None | Full team access via Tailscale mesh network, controlled via ACL |
| Compliance tooling | AWS Security Hub | Falco + Kyverno (simpler surface, fewer tools needed) |

## Support

Direct Slack access day-to-day and fortnightly calls to coordinate priorities. We respond to alerts, handle capacity planning, and help debug application-level issues — feeding back infrastructure-level findings to their developers with specific remediations.

Their team uses the Grafana observability stack directly. We build dashboards tailored to their needs, so their developers can see how their application behaves across the cluster without needing to be intimately familiar with the infrastructure underneath.

After migrating their workloads on, we observed early batch runs and saw that resource use was non-optimal. High context-switch rates showed the scheduler struggling with bursty CPU demand from individual batch processes. We built custom metrics to measure actual throughput against allocated resources. We constrained CPU allocations to slow individual processes down, making scheduling more deterministic so batch jobs could be packed tightly into the available hardware. We tightened memory limits so that leaking processes would be killed before their usage ballooned and sat idle. Over several days, we continually tuned these limits until cluster throughput was optimised for their workload.

> "Thank you very much for being so responsive, by the way. It is awesome." - Tim McEwan, CTO, XDI Systems


---

# PrepBusiness: Eliminating Mysterious Downtime While Doubling Hardware

**Industry**: SaaS — Amazon Prep & Logistics | **Location**: Canada (Global Customers) | **Website**: prepbusiness.com

> "We couldn't find anyone else offering the combination of engineering time, on-call SLA, and infrastructure that Lithus provides. Instead of just migrating our infrastructure, they solved problems we'd been struggling with for months, without us having to hire additional staff." - Keith Brink, Founder, PrepBusiness

## The Situation

PrepBusiness is a Canadian SaaS company that provides a warehouse management system for Amazon prep centres. Their customers are global, so uptime matters.

They were running on AWS with a mix of ECS, EKS, and RDS. The setup worked, but sporadic outages kept hitting customer-facing services. Nobody could figure out why. The RDS instance was under-resourced, they were relying on external contractors for AWS maintenance, and they had no visibility into what was actually happening inside their stack.

## What We Built

**Database migration**: Secure network connections between their AWS infrastructure and our cluster, then live database synchronisation with a new redundant PostgreSQL cluster on StackGres. PrepBusiness opted for a weekend maintenance window rather than a more involved zero-downtime switchover. The cutover itself took 15 minutes, well within the maintenance window PrepBusiness had communicated to their customers.

**Private cloud cluster**: A single-AZ cluster running our [full managed Kubernetes stack](/services) that more than doubled their available compute and memory at the same monthly cost.

**KubeVirt for VM workloads**: KubeVirt for workloads that need full VMs. This includes dedicated GitHub Actions runners that we operate on PrepBusiness's behalf, so their team gets CI/CD runners without managing the infrastructure.

## Finding the Downtime

The mysterious outages were one of PrepBusiness's main pain points, and they persisted after migration. That told us the problem was not infrastructure-related.

We used our [observability stack](/services#observability) — Prometheus metrics, Loki logging rules, and Grafana dashboards — to instrument the database layer over several weeks. The data eventually showed what was happening: long-running application queries, PostgreSQL autovacuum, and scheduled maintenance jobs were all landing on the database at the same time. CPU usage would peg at around 20 cores for extended periods. With the database saturated, web requests that depended on it would back up and time out after 30 seconds.

During the investigation, we used the spare capacity in the cluster to double the database's available CPU and memory at no extra cost. Once we could see the overlapping loads in the dashboards, we rescheduled the maintenance jobs and tuned the query patterns so they no longer competed with production traffic. The downtime stopped.

> "We were struggling to pin down an issue that would cause unpredictable downtime for all our customers. Lithus dug into it with us, built dashboards to help us observe performance across all our services, and ultimately isolated the issue so that we could push out a fix." - Keith Brink, Founder, PrepBusiness

## The Results

- **Doubled compute and memory** at the same monthly spend
- **Downtime eliminated** by tracing overlapping database loads to their source
- **45-day migration** from contract to go-live, with 15 minutes of planned downtime
- **No more relying on external contractors** for day-to-day infrastructure work

## Ongoing Work

Since the migration, we have upgraded PrepBusiness's deployment pipeline from NGINX Unit (now deprecated) to FrankenPHP, handled PHP version upgrades, and applied security patches when CVEs landed. We continue to operate their cluster and respond to infrastructure issues so their team can focus on the product. This is what [included SRE time](/pricing) looks like in practice.


---

# Futurepump: IoT Telemetry From Google Cloud to Stable Bare Metal

**Industry**: IoT — Solar Water Pumps | **Location**: United Kingdom | **Website**: futurepump.com

> "As a small company, having a team we can call on who know how our infrastructure is set up is really valuable. I would recommend Lithus to anyone who's looking for a reliable and knowledgeable team." — Richard Stirzaker, Futurepump

## The Situation

Futurepump builds solar-powered water pumps for farmers in developing countries. Tens of thousands of pumps in the field, each sending telemetry data back to a central pipeline for monitoring and diagnostics.

Their data processing system was running on Google Cloud. Costs were growing at roughly 20% per year with frequent spikes, and nobody on the team had the infrastructure expertise to understand why or optimise the setup. The system worked, but it was expensive, opaque, and fragile.

## What We Built

We migrated the entire data pipeline to a bare-metal cluster running our [managed Kubernetes stack](/services). The new stack receives IoT telemetry via an MQTT broker, processes it through the pipeline, and persists it into PostgreSQL with TimescaleDB for time-series data.

We deployed a full observability stack (Prometheus, Loki, and Grafana) so Futurepump's team has visibility into the pipeline that they never had on Google Cloud. When something goes wrong, they can see what is happening and why.

Moving from shared cloud instances to dedicated bare-metal hardware cut request latency by 50%. That improvement landed as soon as traffic hit the new cluster. Dedicated hardware means no noisy neighbours, no shared storage contention, and consistent performance.

## The Results

- **50% reduction in request latency** from moving to dedicated hardware
- **Costs stable for over two years** after growing ~20%/year on Google Cloud
- **Pipeline visibility** through Prometheus, Loki, and Grafana

Futurepump has been a client since before Lithus was formally established. They no longer spend time managing their data pipeline, and we continue to support their infrastructure as it grows. See our [pricing](/pricing) to understand what this kind of engagement costs.

