AI consulting
Potential analysis, cost-benefit, roadmap – independent recommendations without product ties.
- Potential & maturity analysis
- Use-case assessment
- Roadmap & concept
We take an AI project through all three stages: we advise vendor-independently on what is worth doing, we build the system with open-weights models, and we keep it running – on your premises or hosted in Europe. You decide where the data lives; we document what that choice means for cost, operation and data protection.
We offer three things: vendor-independent consulting – potential and maturity analysis, use-case assessment, roadmap – on-premise AI with GPU server hardware installed at your site and open-weights LLMs running locally, and finished solutions with training, from RAG knowledge assistants to prompt engineering. Each can be commissioned on its own, and every recommendation stays free of product ties.
Potential analysis, cost-benefit, roadmap – independent recommendations without product ties.
Local AI systems, set up on your premises – data stays entirely in-house.
Concrete applications, prompt engineering and team enablement – so AI actually delivers.
We compare three ways to run AI so the decision is made on facts, not on hype. US cloud AI is quick to start but moves data out of Europe; our hosted option keeps it in European data centres at a predictable cost; on-premise keeps everything in your building, with air-gapped operation possible, against an upfront hardware investment.
| Criterion | US cloud AI | nokkela.ai · hosted | On-premise (on your premises) |
|---|---|---|---|
| Operation / location | US cloud, worldwide | Data centres in Europe | On your premises |
| Data sovereignty | Data leaves Europe | Data stays in Europe | Data stays in-house |
| GDPR | Third-country transfer critical | GDPR-compliant (EU) | No external processing |
| Cost model | Per token / subscription | Predictable (per token / package) | Upfront investment + running costs |
| Vendor lock-in | High | Open (Nokkela fine-tuned open weights & various third-party models) | Open, full control |
| Offline operation | No | No (hosted) | Air-gapped possible |
We start where the benefit is measurable today: searchable internal knowledge through RAG, automatic capture of invoices, contracts and emails, support and FAQ assistance, analyses and reports, drafted texts with quality control, and workflow automation together with nokkela.io. We assess each use case for effort and value before anything is built, so nothing is automated for its own sake.
Make internal documents searchable (RAG) – answers instead of searching.
Automatically capture and structure invoices, contracts and emails.
Assistance for support, FAQ bots and draft texts.
Detect patterns, create reports, support decisions.
Marketing and technical texts efficiently, with quality control.
Streamline workflows – often together with nokkela.io.
We choose the model for the task, not for a vendor relationship: Mistral, Qwen, Gemma 4, GPT-OSS, Nemotron and DeepSeek all run locally and GPU-accelerated. Speech models such as Whisper and Kokoro cover speech-to-text and speech synthesis. Because the weights are open, the model stays runnable on your own hardware and you are not tied to a single provider.
Below are the questions we are asked most often about this area. Each answer is written to stand on its own, so it stays correct when quoted alone.
Open-weights models are models whose trained parameters are published for download, so nokkela.ai can run them on hardware you control instead of calling someone else's API. That is what makes the on-premise and European-hosted options possible at all: the model file sits next to your data. Licence terms differ from model to model, so we check each one for your intended use and name it in the concept.
nokkela.ai works with a broad catalogue of open-weights models – among them Mistral, Qwen, Gemma 4, GPT-OSS, Nemotron, DeepSeek, GLM, MiMo, MiniMax, Kimi-K2 and Phi – plus speech models such as Whisper for speech-to-text and Kokoro for speech synthesis. The choice is made vendor-independently: use case, the languages involved, the answer quality required and the available GPU memory decide it, and we test the shortlist against your own material before committing.
An on-premise AI system from nokkela.ai needs a GPU server, a local inference service and integration into your existing systems; we specify, procure and install the hardware on your premises. Sizing follows the workload: the size of the chosen model and the number of concurrent users determine how much GPU memory is needed, so we fix that in the concept rather than guessing. Network, storage and backup are planned with it.
An on-premise AI system from nokkela.ai can be operated air-gapped, because the open-weights model and the retrieval index sit on your own hardware and no request goes to an external API. Model files, the inference stack and updates come in through your own change process instead of a live download. The hosted variant cannot do this – it runs in data centres in Europe and needs a connection.
nokkela.ai offers both: hosted means the models run in data centres in Europe, on-premise means they run on your hardware, in your building. Hosted keeps the cost predictable – per token or as a package – and needs no hardware investment from you; on-premise means an upfront investment plus running costs, but the data never leaves your premises and offline operation is possible. Both avoid a third-country transfer to a US cloud.
A RAG knowledge assistant from nokkela.ai first searches your own documents and then has the language model answer on the basis of the passages it retrieved, so staff get answers instead of hit lists. What we need is access to the document sources – file shares, wikis, ticket systems, contracts – plus your access rules, because the assistant must respect the permissions that already exist. We start with one clearly bounded body of documents rather than everything at once.
With nokkela.ai, whatever you put into an on-premise system stays inside your network: the model runs on your own hardware, so prompts, documents and the index built from them are not sent to an external provider and are not available for anyone else's training. In the hosted variant, processing takes place on infrastructure in Europe instead of with a US provider. Where a use case calls for a third-party model, we name it and its terms in the concept beforehand.
We start with a free initial consultation of up to 30 minutes: you describe the use case, we say honestly whether AI is the right tool and which path – hosted or on-premise – fits. We usually reply to enquiries within one business day. A written concept with scope and cost follows before anything is implemented.
nokkela.ai – a service of Nokkela-IT-Concept GmbH
Email: [email protected]
Need AI agents? nokkela.io