HostYourAI offers four ways to run open models, side by side on a single account and credit balance: shared per token, dedicated GPUs per minute, private single-tenant with a VPN, and self-hosted on your own hardware. Start with EU Hosted inference and move to dedicated capacity when the workload or compliance profile requires it.
1. EU Hosted Gateway: pay per token
One OpenAI-compatible API key, model catalog, hyai/auto, and scale-to-zero shared capacity. Best for SaaS integrations, agencies, agent apps, and experimentation. De tarieven per model (EUR per miljoen tokens), exact de prijzen waarmee je verbruik wordt afgerekend:
De actuele modeltarieven staan in de catalogus in je dashboard; vraag ze anders op via /contact.
EU Hosted means the request is processed in the EU or the wider EEA. The shared Router serves from EU-established providers first, and any machine we rent qualifies only if its data centre sits in that area, with no exception for price or availability. EU Sovereignty Mode goes a step further and restricts processing to sub-processors that are themselves established in the EU, with audit export and support-access controls on the account. Both are set out party by party on our sub-processors page.
2. Dedicated EU Deployment: billed per minute
You pick a GPU class and region, deploy your own vLLM instance, and pay for as long as it runs. Best for custom Hugging Face models, BYOK upstreams, steady high-volume workloads, or when you need full control over the deployment.
The prices below are hourly rates, but you are billed per minute, with no rounding up to the full hour. Delete an instance after six minutes and you pay for six minutes.
GPU pricing follows live EU availability at our providers, so we quote it per deployment rather than from a fixed table. The exact hourly price is shown before you deploy. Ask us for a class we should keep warm for you.
Liever met je echte cijfers? Op de vergelijkpagina plak je een OpenAI- of Anthropic-adminsleutel, vul je je tokens in of kies je je ChatGPT- of Claude-abonnement, en zie je wat hetzelfde verkeer hier kost. Zonder account, er wordt niets bewaard. Vergelijk je kosten
3. Private single-tenant: dedicated with VPN, on request
Need an isolated runtime with dedicated GPUs per customer, at-rest encryption, and a private network policy? This is a dedicated deployment on our own EU hardware where the node itself is the WireGuard endpoint: you talk to the model through the tunnel and the model port is closed to the public internet. For healthcare, government, legal, finance, and workloads that cannot use shared capacity, we scope and price this per project, usually as a fixed monthly fee without token charges. The configurations below are typical starting points, not a self-serve product.
| Configuration | VRAM | Indicative / month | Setup (one-off) |
|---|---|---|---|
| 1× L40S | 48 GB | from € 1,200 | € 500 |
| 1× H100 | 80 GB | from € 3,500 | € 1,000 |
| 2× H100 | 160 GB | from € 6,500 | € 1,000 |
| 4× H100 | 320 GB | from € 12,500 | € 1,500 |
Indicative, scoped per project. Talk to us via /contact. Confidential computing (TEE) is on the roadmap; we will not price what we have not yet validated.
4. Self-hosted: the Orchestrator on your hardware, on request
Some organisations may not have any external party in the path of a request, or already own GPUs in their own data centre or cloud account. For them we install the Orchestrator, the software that runs our own platform, on their servers: model garden, deployment, the Router with key management and usage reporting, and the health checks that bring a fallen node back. We set it up, hand over keys and operations, and remove our access; updates come as releases and operations by us are optional. Priced as an annual licence plus a one-off setup, scoped per project via /contact.
BYOK: bring your own API key
You can attach your own OpenAI, Anthropic, Google or Mistral API key to an instance. We forward your traffic to the upstream under your contract with them. BYOK currently carries no platform fee: you only pay your own provider. Useful for hybrid setups that mix EU-hosted open-weights with frontier closed models. Per API key you can also let Autopilot (hyai/auto) pick those own-key models for heavy requests; light and medium work stays on open EU models, and the switch does not exist on an EU-only key.
Getting started
- Creating an account is free. No credit card to sign up.
- Pay as you go from a single prepaid credit balance. No subscription, no minimum.
- Top up with iDEAL, card or SEPA, then call the Router or deploy an instance.
Billing
- Currency: EUR. VAT added where applicable; reverse-charge for EU B2B with valid VAT number.
- Method: Stripe (credit card, iDEAL, SEPA direct debit). Invoices auto-issued from your dashboard.
- Credits: top up in advance; balance is consumed by all hosted modes from a single pool.
- Volume / partner tier: for €> 2 000 / month in tokens or one or more single-tenant deployments, we offer a partner tier with discounts, SLAs, and a dedicated technical contact. Contact info@hostyourai.com.
What's not on the price list
Bespoke procurement, custom contracts, NEN 7510 / BIO audit packages, white-label / reseller arrangements, and confidential-computing deployments are quoted per project. Talk to us via /contact.