Your work.
Your models.
Your terms.
Cost-effective physical hardware dedicated to your organisation alone. Never shared with another customer.
Discuss your needsWe bring together dedicated hardware, carefully selected models and the expertise to make them work well. You decide where it runs, how your data is handled and which capabilities your team needs.
- Analyse and process confidential or proprietary data
- Deploy sovereign agentic workflows
- Generate and edit media showcasing classified or unreleased products
Keep the decisions that matter.
Agree the operating arrangements that matter to your team: data policies, model capabilities and controlled changes.
Your data and policies
You own your data and decide audit logging, retention and usage policies. Your data never leave your dedicated hosts, so Kalufs does not inspect your data unless you let us.
Why European soil is no guarantee against foreign interests acquiring your data →
Why privacy is not the whole confidentiality question →
Models that can do your work
Legitimate work can involve sensitive material. We help you select and validate models, including reduced-refusal or abliterated variants, when your workflows need them. You set the usage policies.
When safeguards block legitimate work →
Changes made deliberately
We agree model changes with you and validate them against your workloads. Kalufs configures and tunes the agreed models to make effective use of the hardware.
Why model lifecycle is part of the service →
Different capabilities.
A shared way in.
Your applications request an agreed model or role alias. We configure what stays ready and what loads on demand with you.
Run models from families such as DeepSeek, Gemma, GLM, Ideogram, Kimi, Krea, LTX, MiMo, Minimax, Mistral, Muse, Qwen and YuE on your own terms.
An engineering team
- GCD 0Reasoning model · shard
- GCD 1Reasoning model · shard
- GCD 2Reasoning model · shard
- GCD 3Reasoning model · shard
- GCD 4Prose model · shard
- GCD 5Prose model · shard
- GCD 6Image model XImage model Y
- GCD 73D asset generation modelOCR model
Model loading, host workloads and optional KV-cache offload share this capacity.
Reasoning model, prose model, 3D asset generation and OCR remain resident in this example. Image model X and Image model Y are exposed as separate model endpoints; the selected image model is swapped into GCD 6 on demand.
A research and creative studio
- GCD 0Prose model · shard
- GCD 1Prose model · shard
- GCD 2Prose model · shard
- GCD 3Prose model · shard
- GCD 4Video generation model · shard
- GCD 5Music generation model
- GCD 6Image model XImage model Y
- GCD 7Text-to-speech model3D asset generation model
Model loading, host workloads and optional KV-cache offload share this capacity.
Prose model, video generation model, music generation model, text-to-speech model and 3D asset generation model remain resident in this example. Image model X and Image model Y are exposed as separate model endpoints; the selected image model is swapped into GCD 6 on demand.
A legal team
- GCD 0Prose model · shard
- GCD 1Prose model · shard
- GCD 2Prose model · shard
- GCD 3Prose model · shard
- GCD 4Image analysis modelVideo analysis model
- GCD 5Custom legal research + citations model
- GCD 6OCR modelSpeech transcription model
- GCD 7Translation modelStructured extraction model
Model loading, host workloads and optional KV-cache offload share this capacity.
Prose model, custom legal research + citations model, OCR model, speech transcription model, translation model and structured extraction model remain resident in this example. Image analysis model and Video analysis model are exposed as separate model endpoints; the selected analysis model is swapped into GCD 4 on demand.
Each tile is one GPU compute die (GCD), not a whole graphics card or a virtual security boundary. Repeated colours show one model distributed across multiple GCDs. Occupied tiles identify allocations. Model names and labels are illustrative.
Diffusion models can share a GCD when their combined weights and execution memory fit. Typical diffusion generation does not need the growing autoregressive KV cache used by an LLM, but activations, attention buffers and image/video decoding still need memory.
Supported serving stacks can offload LLM KV cache to system RAM. RAM remains on the same dedicated host.
Exact models, memory use and loading arrangements are agreed for your work.
A stable name. An agreed upgrade.
As new, better models are released or your needs change, we can replace models at your request. We ensure they run efficiently on your dedicated hardware and can help customise, fine-tune and adapt foundation models to your needs. In every case, we provide a stable, agreed upgrade path.
Start with one node.
Expand with your work.
Each node has ConnectX-6 200GbE connectivity. Add dedicated nodes as demand grows: to serve more work in parallel, or to accommodate a larger model across nodes with a compatible serving stack.
We agree and validate node count, model partitioning, network topology and performance for your workload. Additional nodes follow the same agreed operating arrangements.
Close to your work.
Inside your boundary.
Choose your deployment
Choose an on-premises deployment or hosting in Östersund (Sweden). Hosting in Imperia (Italy) is coming soon! Location is a separate choice from which models stay loaded.
On your premises
Dedicated hardware is physically installed at your premises. You can choose an air-gapped deployment for extra isolation.
Updates and modifications require physical access to the host. This can lengthen support response and resolution times. Maintenance arrangements and support expectations are agreed as part of the deployment.
Östersund
In the middle of Sweden, not far from Trondheim in Norway and near the mountains, Östersund is “Vinterstaden”—“the Winter City”. Our partner’s data centre is supplied with fossil-free hydropower.
Imperia
Coming soon! Our partner’s data centre in Imperia is under construction and will soon be operational.
On the Ligurian coast at the foot of the Ligurian Alps, less than an hour from the Nice metropolitan area in France, Imperia celebrates its slogan “il miglior clima d’Italia”—“Italy’s best climate”. Free cooling reduces cooling energy demand, while abundant sunshine supports fossil-free solar generation.
More than a host.
A reliable partner.
We help you put the hardware to work, from deployment and integration to ongoing support.
Kalufs offers consultancy in software development, agentic workflows, systems integration, data labelling, training operations, MLOps, red teaming, penetration testing, security monitoring, incident response and managed fine-tuning.
Start with the work you need to do.
A useful first conversation covers:
- Your area of business, your data-sovereignty challenges and opportunities.
- Your workflows and required model capabilities.
- Where processing should happen and who may access the data.
- The integrations, evaluations and support your team needs.
Interested? Send us an e-mail at info@kalufs.se for an informal conversation about your needs.
Different reasons
to take control.
Articles on enduring questions, updated in place as the topics develop.
- Sovereignty and jurisdiction
US interests can reach your data inside a European data centre
Why hosting in Europe does not by itself remove US jurisdiction from your provider, and what foreign-intelligence law covers.
- Confidentiality
Privacy is not the whole question
Personal information, valuable knowledge and the gap between a promise and the ability to verify it.
- Capability
When safeguards block legitimate work
What Hugging Face’s incident response tells us about having a capable local model ready.
- Lifecycle
The same model name can behave differently
Why configuration, evaluation and controlled change matter alongside model quality.
Your privacy
We value your privacy. We use no analytics cookies and do not track you across websites. We use Plausible Analytics only to understand aggregate website traffic, such as visitor numbers and general geographic locations. We respect Do Not Track and Global Privacy Control.