Enterprise AI is bottlenecked by infrastructure. Organizations are integrating AI into their value chains. To protect IP and sensitive data, doing this in a sovereign way (retaining full-stack ownership rather than relying on black-box APIs) is an absolute necessity. However, the underlying compute resources are typically fragmented across private clouds and mixed generations of on-prem hardware. Utilizing these expensive GPUs efficiently is the primary ROI lever. This is fundamentally an infrastructure orchestration challenge. Here is how exalsius is solving this: ✅ Maximizing ROI on Expensive Assets: GPUs are costly, and their ROI is strictly tied to utilization rates. By pooling isolated GPU capacity into a unified compute system, we ensure these highly valuable assets do not sit idle. ✅ Enabling Geo-Distributed Resource Sharing: GPU resources often bottleneck in one region while sitting unused in another. We enable seamless resource sharing across geographic boundaries. For example, we successfully bridged Nvidia A100 nodes in Canada with AMD MI300X nodes in the US. This cross-border setup is managed cohesively by a central control plane in Germany. ✅ Centralizing Cost Control: Managing compute that is distributed across multiple clouds and on-prem sites is difficult. Yet, it is essential today due to ongoing GPU supply shortages. Our platform presents a consistent operational interface, regardless of whether the underlying infrastructure relies on AWS, Nebius, bare metal, or other providers. Together with Mirantis, we published a technical breakdown of our architecture on the official Cloud Native Computing Foundation (CNCF) blog. The post presents our findings on executing AI workloads across cross-vendor GPUs, managing cross-site networking, and implementing dynamic, energy-aware AI orchestration. (Link to the technical post in the comments) Hardware heterogeneity is the new normal. Extracting real business value from AI requires mastering distributed infrastructure orchestration. This work is funded by SPRIND - Bundesagentur für Sprunginnovationen. Randy Bias Prithvi Raj Jussi Nummelin Soeren Becker Dominik Scheinert Philipp Wiesner
exalsius
Technology, Information and Internet
Berlin, Berlin 293 followers
Providing AI Teams with One-Click GPU Infrastructure
About us
exalsius simplifies access to compute infrastructure for AI teams. Training today’s AI models often demands multiple GPU nodes — either because the models are too large for a single GPU, or the data volumes require distributed processing to finish training within a reasonable timeframe. Therefore, building AI is no longer just about designing smart models and training algorithms. It’s become an infrastructure problem — one that forces teams to search across cloud providers for affordable GPUs, configure distributed servers and networking, install and maintain complex dependencies, and struggle to keep jobs running reliably and securely. We believe AI development should focus on innovation — not infrastructure. That’s why we’re building **exalsius**: a Kubernetes-native system that automatically finds affordable GPUs across clouds or data centers, prepares the infrastructure, installs the right tools, and deploys your distributed training jobs — securely, reliably, and without hassle. Whether teams are using Ray, kubeflow, or custom pipelines, exalsius handles everything beneath — provisioning, setup, and orchestration of the GPU nodes — so they can concentrate on what matters most: building great AI.
- Website
-
https://exalsius.ai
External link for exalsius
- Industry
- Technology, Information and Internet
- Company size
- 2-10 employees
- Headquarters
- Berlin, Berlin
- Type
- Privately Held
- Founded
- 2025
- Specialties
- GPU Infrastructure, Model Deployment, AI Training, AI Inference, and Distributed Training
Locations
-
Primary
Get directions
Max-Urich Straße
Berlin, Berlin 13355, DE
Employees at exalsius
Updates
-
exalsius reposted this
We have enough data and great algorithms. Infrastructure is the main AI bottleneck. There is a huge amount of high-quality data available, and great models are openly accessible. Companies want to build on this. They want to deploy models, integrate them into their value chains, automate processes, and fine-tune them to their specific needs. The ideal setup for this is a homogeneous, centralized data center or a cloud with infinite resources to host the AI. The reality is very different. Compute resources are highly fragmented across different clouds or internally on different machines. We are dealing with different GPUs, slow and fast interconnects, and conflicting setups. It is messy. This is fundamentally a cloud-native infrastructure problem. We are building a solution to orchestrate AI training and inference across this fragmented infrastructure exalsius. In our recent project, we are running AI on Nvidia GPUs in Canada and Germany, connected with AMD GPUs in the US. To share how we did it, we teamed up with Mirantis and published a technical deep dive on the official Cloud Native Computing Foundation (CNCF) blog. We cover the challenges and lessons learned around AMD and Nvidia GPU orchestration, cross-site networking, and dynamic energy-aware AI scheduling, and more. (Link to the blog post in the comments) Hardware heterogeneity will only increase. Especially in Europe, our infrastructure will remain fragmented and distributed. Getting real value from AI is and will increasingly be an infrastructure orchestration challenge. This work is funded by SPRIND - Bundesagentur für Sprunginnovationen Prithvi Raj Randy Bias Jussi Nummelin Soeren Becker Dominik Scheinert Philipp Wiesner
-
-
exalsius reposted this
I had a great time at KCD Czech & Slovak last week, both attending and speaking. Together with Jussi Nummelin, I gave a talk on "Breaking Free of a Single Datacenter: Expanding Geo-Distributed AI Platforms". The talk was about moving AI infrastructure beyond the single-datacenter assumption: fragmented GPU resources, mixed hardware, geo-distributed sites, and what it takes to make this operational with Kubernetes, k0s, k0smotron, and k0rdent. Part of the talk focused on field studies we conducted with exalsius, including geo-distributed training across heterogeneous GPUs and energy-aware training during curtailment windows. Thanks to everyone who attended, asked questions, or stopped by afterwards. I really enjoyed the talks and conversations at the event. #kcdczsk26 #K0s #K0rdent #K0smotron SPRIND - Bundesagentur für Sprunginnovationen
-
-
exalsius reposted this
A few weeks ago, I had the opportunity to attend the ICPE 2026 (https://icpe.spec.org/) in Florence, Italy, and present our work on reliability and observability for modern LLM inference systems at the QualITA workshop. As LLMs increasingly power critical applications, ensuring reliable and observable AI infrastructure is becoming an important systems engineering challenge. At the same time, the field is evolving incredibly fast: architectures, serving stacks, and operational best practices are continuously changing, with many promising developments emerging across both industry and academia. In our paper, “Beyond Microservices: Testing Web-Scale RCA Methods on GPU-Driven LLM Workloads”, we investigated how existing root cause analysis (RCA) approaches, originally designed for traditional microservices, generalize to modern GPU-driven LLM inference environments. Using controlled fault injection on a production-style serving stack, we evaluated a broad range of RCA methods across CPU, memory, network, and GPU-related anomalies. While our reference architecture and findings should be interpreted in the context of a rapidly evolving ecosystem, we observed both important limitations of existing approaches and many encouraging signals from the community regarding observability, inference infrastructure, and AI operations tooling. This work represents one research artifact from a concluded joint research collaboration between IT-systems engineers from logsight, Technische Universität Berlin, and Technology Innovation Institute. At exalsius, we continue exploring related challenges around cloud-native AI infrastructure, including optimized, scalable, and observable LLM inference deployments as one important use case among others. Publication: https://lnkd.in/evGntcC5 Preprint: https://lnkd.in/evk5TcFP Repository: https://lnkd.in/emxAHDT6
-
-
exalsius reposted this
We are not late to the AI party. We are early to a different, European one. My main takeaway from the Flower Summit 2026 was that we aren’t here to win someone else’s race. We are building an ecosystem that actually fits the way Europe works: our federal structures, our distributed, highly valuable industry data, and our human-centric values. Instead of trying to force-fit a centralized, homogeneous model onto a decentralized continent, we are building a new AI paradigm designed for heterogeneity and federations. Proud to be part of this accelerating ecosystem and this new way of thinking about AI. Thanks for the great event, Flower Labs, Daniel, Nicholas, Taner! 🌼
-
-
exalsius reposted this
Getting Germany’s entire deep tech ecosystem into one room. That’s something only SPRIND can pull off. This year’s Venture SPRIND felt different. Compared to last year, there was a new level of urgency. A pressure to move, to build, to race paired with a deep sense of optimism. I had the chance to sit down with Johannes and Arya to discuss a critical tension: the mismatch between the current AI ecosystem and Europe’s federal, data-protective structure. Building at the forefront of AI we are faced with a choice: 1. We can try to copy the American or Chinese approaches to AI. 2. Or, we can build AI that actually fits our European structures and values. Both are tremendously hard challenges. Personally, I’m in strong favor of the second one. That is why we are building the fundamental infrastructure for federated, distributed, and decentralized AI. This means building AI that works for us, not against our principles. Thanks for support us on that journey, and thanks for having me SPRIND - Bundesagentur für Sprunginnovationen, Johannes, Jano, Leonard, Marcia, Sebastian.
-
-
exalsius reposted this
Last week we published a technical report on distributed LLM pretraining during renewable curtailment windows, now available on arXiv (link in comments), which demonstrates the awesome potential of GPU resource flexibility and intelligent workload scheduling. The core result: we trained a 561M-parameter language model across GPU clusters in California, Texas, and South Australia, activating training only when regional renewable generation exceeded grid demand. The system elastically switches between local and federated training as sites come online or drop off. Training quality matched conventional single-site runs while producing just 5-12% of the carbon emissions. 🌱 I wrote a companion piece on Medium (link also in comments) exploring an application of this architecture beyond language models, specifically for multi-agent reinforcement learning (MARL) in distributed energy systems. A few key points from that piece: - MARL training for energy systems shares properties that make LLM pretraining a good fit for curtailment-aware scheduling: it can be delay-tolerant, checkpointable, and distributable across sites. - The federated structure we demonstrated (elastic participation, privacy-preserving weight sharing, geographic distribution) maps directly onto how MARL agents for virtual power plants would need to train. - GPU clusters running these workloads can themselves participate in grid flexibility, creating a closed loop where the training infrastructure is also a participant in the system being optimized. - For smaller, purpose-built clusters designed around delay-tolerant workloads and strategically located near stranded energy, the economics of grid flexibility can be more attractive than for hyperscale facilities that are optimized for sustained high utilization. Open questions remain around whether federated averaging preserves learning quality for interacting agent populations, and how sensitive MARL convergence is to the kind of intermittent training that curtailment windows impose. I'll be working on answering these questions during the rest of my time in the Venture Science Doctorate and beyond, but for now, the systems infrastructure exists and the application domain is well-matched. Special thanks to Philipp Wiesner, Alexander Acker, and the rest of the exalsius team for their expertise in refining and executing on the project's goals. Thanks to WattTime.org for the use of their API and to Flower Labs for their federated training infrastructure. And thanks to Deep Science Ventures for everything else.
-
-
exalsius reposted this
I spent the weekend building `autoimprove` (https://lnkd.in/d6RA-g8g), a package that brings autoresearch improvement loop to any repository. The idea is creating autoresearchable (improvable) version of any repo, ML and non-ML. How to use autoimprove: - clone the repo - open claude code / opencode or similar - tell it: `install autoimprove cli and run it on my repo "<repo_path>"` - inspect `.autoimprove/program.md` - let it run I tried it on my own side project`opentab`: https://lnkd.in/dqSpzbDC On `opentab`, `autoimprove` proposed and tested changes like: vectorizing `DecisionTreeMapping`, pre-norm instead of postnorm, better truncation / test target handling, improving test embedding pooling, adding label smoothing. After several iterations (tried more than 30), it moved the average score on baseline evaluation datasets from 44% to 91% in just 4 minutes of training on laptop gpu. Fixed a lot of small errors. My take is that in the research space, these loops will provide speed up most definetely. However, having clear idea what to do still requires strong knowledge of fundamentals and first principles. I still can't speak for general software. If you want to try it, break it, improve it, or contribute directly, I'd really love that: https://lnkd.in/d6RA-g8g thanks to exalsius for providing GPUs for testing.
-
-
Resolving the paradox of AI scaling vs. sustainability Great to see our team member (Congratulations, Philipp!) on stage at the Flower AI Summit 2026 in London. Large-scale AI training is notorious for its energy demands, but Philipp is proving it doesn't have to stay that way. His upcoming talk, “Toward Carbon-Aware Distributed LLM Pretraining,” dives into a future where training jobs are mobile, shifting across time and geography to sync with periods of cleaner, greener electricity. Join the Flower Labs, us, and a huge community of distributed AI folks in London (April 15-16) to discuss a different approach to AI scaling.
-
-
exalsius reposted this
The SPRIND Composite Learning Challenge has shifted from "Can we build it?" to "How can Europe win AI?" Most conversations in the EU ecosystem revolve around how to avoid losing. I'd rather think about how we can win AI. We have unique structural advantages: -> The European industry owns a huge amount of data that far outweigh what current LLMs are trained on. -> Our fragmented, diverse SME landscape is actually a much better field for building agentic models. Diverse organizational structures allow models to generalize better than on large monolithic US corporate structures. Europe's GPU infrastructure is fragmented and scarce. No single entity has the scale of a US hyperscaler. Jointly pooling resources is the only viable path forward. All Composite Learning Challenge teams proofed that the technology is working. The challenge now is to solve Europe's "federated friction." Reaching consent across a multitude of stakeholders is slow. In a space that moves as fast as AI, this hurts. We are spending a lot of our time figuring out how to make it undeniably beneficial for organizations to pool data and compute and how to implement a structure that maintains IP ownership. Scaling Composite Learning is a massive challenge by definition, and we need a fundamental shift in how European organizations collaborate. Flower Labs, Daniel J. Beutel, Taner Topal PanocularAI, Arya Mazaheri, Sören Heß SPRIND - Bundesagentur für Sprunginnovationen, Jano Costard, Johannes Otterbach, Leonard Schenk, Marcia Holst
-