Management API

GPU instances, API keys, teams and their routers, scriptable with your regular API key.

Use the Management API to script the full lifecycle of your resources: create a dedicated GPU instance, use it, delete it when you are done, and manage your API keys (create, rotate, revoke). Everything runs on the same host (https://hostyourai.com) and authenticates with the same hyai-rt-* key you use for inference. There is no separate token type.

1. Authentication

Send your API key as a Bearer token with every request:

Authorization: Bearer hyai-rt-...

Create a key in the app under API keys, or deploy your first instance: an account without a key gets one automatically. A key manages only the resources of its own account. Script with a personal key. A key that is shared with a team of more than one person is held by every member of that team, so it is refused (403 personal_key_required) as soon as that team has a router, and this is being extended to every shared team key. Embedded tenant keys cannot manage anything. Teams and routers always need a personal key.

2. Instances

An instance either exists or it does not. There is no pause state: while it runs you pay the hourly price per minute, and deleting it stops billing immediately. Recreating one later is just another create call.

List

GET /api/v1/instances

Create (deploy)

curl -X POST https://hostyourai.com/api/v1/instances \
  -H "Authorization: Bearer hyai-rt-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model_id":      "llama-3.2-1b",
    "region":        "EU",
    "instance_type": "text"
  }'

Same fields, validation and credit check as the deploy wizard. Optional: custom_model_path (any HuggingFace repo), custom_min_vram, gpu_name, gpu_count, hf_token (your own HF token for gated repos, encrypted at rest). Returns the instance with status: "deploying"; a background job allocates the GPU and pulls the model.

Status

GET /api/v1/instances/{id}

Status values: pending, deploying, running, failed. Poll at most once per few seconds. While your own instance is still deploying, inference calls for its model return 503 with a Retry-After header and an honest ETA.

Delete

DELETE /api/v1/instances/{id}

Destroys the GPU at the provider; billing stops immediately. If the provider does not confirm the teardown, the instance stays visible so nothing keeps billing invisibly.

Private connection (WireGuard VPN)

Add "vpn": true to the create call and the instance becomes a WireGuard endpoint on our own EU hardware. You talk to the model through the tunnel, directly on the GPU, without our platform in between. The model port is closed to the public internet; only the tunnel and our platform (health checks, Playground) can reach it. Optional: vpn_client_public_key (your own WireGuard public key, base64). Without it we generate the client key pair for you.

curl -X POST https://hostyourai.com/api/v1/instances \
  -H "Authorization: Bearer hyai-rt-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model_id": "llama-3.2-1b",
    "region":   "EU",
    "vpn":      true
  }'

GET /api/v1/instances/{id}/vpn

The VPN call returns endpoint (public IP and UDP port of the instance), server_public_key, client_address, base_urls (TLS on port 8000 with the self-signed tls_cert, plain HTTP on port 8080 inside the tunnel only), api_key for calls inside the tunnel, and client_config: a ready-to-use wg0.conf. Keys stay the same for the life of the instance; after an automatic restart on new hardware only the endpoint changes, so re-read this call when the tunnel drops. Available on our own hardware only (UpCloud, Scaleway); the provider is chosen for you when you ask for a VPN.

Direct connection to your own box

GET /api/v1/instances/{id}/direct

For your own router (for example hyai-route): how to reach a running dedicated instance of your account directly, without our platform in between. Returns id, model, base_url (the instance endpoint plus /v1), api_key (the key vLLM was started with, or null), tls_cert (the self-signed certificate to pin, PEM, or null), context (the measured context window, or null) and tool_parser. Rows of GET /api/v1/instances carry "direct": true when this call will answer. 409 instance_not_running while the instance has no endpoint yet; 409 instance_private for an instance with a private connection, which you reach through its tunnel instead.

2a. Balance

GET /api/v1/balance

Returns {"object": "balance", "currency": "EUR", "balance": 41.25, "spend_month": 12.8, "auto_topup": false, "as_of": "..."}: your credit balance as the billing page shows it, and this calendar month's spend (router tokens plus GPU minute charges). A key shared with a team of more than one person always gets 403 personal_key_required here: team members never see the account's balance.

3. API keys

List

GET /api/v1/keys

Create

POST /api/v1/keys
{ "name": "production" }

The plaintext key is returned once, in the key field of the response.

Rotate

POST /api/v1/keys/{id}/rotate

Same key row, new secret. The old secret stops working immediately; the response contains the new one. You may rotate the key you are calling with.

Revoke

DELETE /api/v1/keys/{id}

Put a key in a team

POST /api/v1/keys
{ "name": "knowledge-workers", "team_id": 12 }

PATCH /api/v1/keys/{id}
{ "team_id": 12 }        # or { "router_id": 5 }, or { "team_id": null } to make it personal again

A key follows the router of the team it is in, so router_id and team_id mean the same thing here. Usage stays billed to the owner of the key. GET /api/v1/keys, PATCH and a create with a team return team_id and router_id for every key.

3a. Teams and routers

An account can have several teams, each with its own members, keys, budget and router. Manage them with a personal key: a key that is shared with a team is refused (personal_key_required), because every member of that team holds it.

GET    /api/v1/teams
POST   /api/v1/teams                 { "name": "Knowledge workers" }
GET    /api/v1/teams/{id}
PATCH  /api/v1/teams/{id}            { "name": "..." }
DELETE /api/v1/teams/{id}            # refused while the team still has active keys

GET    /api/v1/routers               # every router of the teams you manage
GET    /api/v1/teams/{id}/router
PUT    /api/v1/teams/{id}/router
DELETE /api/v1/teams/{id}/router

PUT replaces the whole configuration. Leave is_active out and the router keeps the state it has (a new router starts active); false pauses it: the setup stays, the keys of the team follow their own settings. An unknown field, step type, policy or preference, a budget or limit that is not a positive number, a repeated step or more than twelve models is refused with 422; nothing you send is dropped in silence.

PUT /api/v1/teams/12/router
{
  "is_active": true,
  "config": {
    "guards": { "eu_only": true, "monthly_budget": 250, "models": [] },
    "policy": "cost",
    "route": [
      { "type": "own_keys", "scope": "heavy", "monthly_token_cap": 400000000 },
      { "type": "model", "slug": "zai-org/GLM-5.2", "use_for": "code", "monthly_budget": 100 },
      { "type": "model", "slug": "claude-opus", "own": true, "use_for": "heavy", "monthly_token_cap": 250000000 },
      { "type": "autopilot" }
    ]
  }
}
  • Guards apply to every request of every key in the team: only sovereign European capacity, one monthly budget for the whole team, a model list, and the policy Autopilot uses when a request names none (cost, balanced, quality). The strictest of key and team wins.
  • Route applies to requests for hyai/auto, tried top to bottom: your own provider contracts, models with the work they are preferred for (all, light, medium, heavy, code) and their own monthly limit, and Autopilot. A model behind your own contract ("own": true) is limited in tokens, an open model in credits. Without an autopilot step a request that fits no step gets model_unavailable; an empty route leaves Autopilot to decide as always.
  • A request that names a model goes to that model, within the guards.
  • Budget used up: 402 team_budget_exceeded, or 402 team_model_budget_exceeded for one model. Budgets are a soft limit: requests already running when the limit is reached still finish.

Error codes

403 personal_key_requiredTeams, routers and team_id on a key are managed with a personal key, not with a key that is shared with a team.
404 not_foundThe team, router or key does not exist, or is not yours to manage.
409 team_has_keysThe team still has active keys. Revoke them or move them out first.
409 default_teamThe team your account started with cannot be deleted.
422 invalid_requestUnknown step type, policy or preference in a router, a model step without a slug, or a missing name.
403 model_not_allowedInference: the model is not on the model list of the team or the key.
403 sovereignty_violationInference: EU only is on and the machine this request landed on is not sovereign.
503 sovereignty_unavailableInference: EU only is on and the model has no sovereign capacity right now.
508 loop_detectedInference: the request already went out over an own provider endpoint that points back at HostYourAI.
402 team_budget_exceeded, 402 team_model_budget_exceededInference: the monthly limit of the team, or of one model in its route, is used up.
503 model_unavailableInference: no step of the route has a model for this request (a route without Autopilot is strict).
503 team_router_unavailableInference: the router of the team could not be read. The request is refused rather than served without its guards; try again.

4. Inference

The same key does inference on the OpenAI-compatible endpoint. Your own dedicated instances show up in GET /api/v1/models marked "owned": true, under the model name you deployed; tokens on your own hardware cost nothing on top of the hourly price.

curl https://hostyourai.com/api/v1/chat/completions \
  -H "Authorization: Bearer hyai-rt-..." \
  -H "Content-Type: application/json" \
  -d '{"model":"llama-3.2-1b","messages":[{"role":"user","content":"Hi"}]}'

5. End-to-end example: batch job

# 1. Create an instance
curl -X POST https://hostyourai.com/api/v1/instances \
  -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
  -d '{"model_id":"llama-3.2-1b","region":"EU","instance_type":"text"}'
# → { "id": 42, "status": "deploying", ... }

# 2. Poll until running
curl https://hostyourai.com/api/v1/instances/42 -H "Authorization: Bearer $KEY"
# → { "status": "running", ... }

# 3. Run your job with the same key
curl https://hostyourai.com/api/v1/chat/completions \
  -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
  -d '{"model":"llama-3.2-1b","messages":[{"role":"user","content":"..."}]}'

# 4. Done? Delete, billing stops now
curl -X DELETE https://hostyourai.com/api/v1/instances/42 \
  -H "Authorization: Bearer $KEY"

6. Billing

Dedicated instances bill per minute at the hourly price shown at deploy time, from the moment the model is ready and answering; searching, booting and installing are free. Shared router models bill per token. Optional automatic top-up (Settings → Billing) refills your balance from a saved card when it drops below your threshold, so long-running jobs never stall on credit.

7. Conventions

  • Errors follow the OpenAI error shape: { "error": { "message", "type", "code" } }.
  • All timestamps are ISO 8601, UTC. Currency: EUR.
  • Model discovery: GET /api/v1/models (with your key) or GET /api/router/catalog in the app.

Questions?

info@hostyourai.com