---
description: "Build a production LLM API endpoint on Verda: put a LiteLLM gateway on a CPU instance in front of two Serverless Container deployments, with per-application keys, rate limits, token accounting, load spread and health-based failover."
revision_date: 01.10.2026
---

# Tutorial: LLM API Gateway with LiteLLM

## What you'll build

In this guide we will build one API endpoint that applications can use to talk to a language
model running on two Verda Serverless Container deployments. The application sends a prompt
to the gateway and the gateway chooses which deployment handles it.

An application calling a model directly needs the deployment's Endpoint Address and its inference key. As
you add applications and deployments you also need a way to control who can call the model,
how much traffic each application can send and what happens when an endpoint becomes
unavailable. The gateway puts those controls in one place.

By the end of the guide:

- Applications call one model name `verda-qwen` without choosing a GPU deployment.
- Each application can have its own API key and request limit.
- You can compare the tokens returned to the client with the usage stored in PostgreSQL.
- Requests are distributed across two model endpoints and you test routing when one
  endpoint is stopped and restored.

We use **LiteLLM** because it understands OpenAI-compatible model requests and provides
virtual keys, token accounting and model-aware routing in the same service. The GPU
deployments generate the text. LiteLLM runs on a CPU instance and manages access to them.

***

## Architecture

The request travels through two layers. LiteLLM checks the application's key and limits, then
sends the request to one of the two Verda endpoints. PostgreSQL keeps the gateway's key
records and usage history. In this walkthrough the gateway listens only on the CPU instance
itself. The application and the gateway share that machine.

<pre style="overflow-x: auto; white-space: pre;">

                         CPU instance
+--------------------------------------------------------------+
|                                                              |
|  Application --- localhost:4000 ---> LiteLLM ---> PostgreSQL |
|                                         |       keys + usage |
+-----------------------------------------|--------------------+
                                          |
                           +--------------+--------------+
             HTTPS + key A |                             | HTTPS + key B
                           v                             v
                 +------------------+          +------------------+
                 |    Endpoint A    |          |    Endpoint B    |
                 |    Serverless    |          |    Serverless    |
                 | Container (GPU)  |          | Container (GPU)  |
                 +------------------+          +------------------+

</pre>

Both deployments serve `Qwen/Qwen2.5-1.5B-Instruct`. The client-facing name `verda-qwen` is an
alias for that pool.

| Component | Runs on | Purpose |
|---|---|---|
| LiteLLM | CPU instance | Client authentication, rate limits and model routing |
| PostgreSQL | The same CPU instance in this walkthrough | Persistent virtual-key records and usage logs |
| Model A and model B | Two Serverless Container deployments | Generate model responses |

**The gateway is a single point of failure.** Two deployments protect against one endpoint
going down but if the gateway instance stops every request fails.

***

## Prerequisites

- Have a Hugging Face [User Access Token](https://huggingface.co/settings/tokens) with
  `READ` permission. The deployments use it to download the model weights.
- Have ready an
  [Inference API Key](https://docs.verda.com/welcome-to-verda/api-credentials/#create-inference-api-keys)
  from **Credentials → Inference API Keys** in the Verda console.

***

## Step 1: Create the two model deployments

The gateway needs two deployments that serve the same model. The configuration in this guide
uses `Qwen/Qwen2.5-1.5B-Instruct`, so create both deployments with the settings below. They
differ only in their name.

In the [Verda cloud console](https://console.verda.com/signin), go to
**Containers → New deployment** and use these settings:

| Setting | Value |
|---|---|
| Deployment name | `llm-gateway-a` for the first deployment, `llm-gateway-b` for the second |
| Compute type | L40S, which this guide was tested on. General Compute (24 GB VRAM) also works: the [vLLM quickstart](https://docs.verda.com/containers/tutorials/deploy-with-vllm-quick/) runs the same model and tag on it |
| Container image | `docker.io/vllm/vllm-openai`, with **Public** toggled on |
| Tag | `v0.29.0` |
| Exposed HTTP port | `8000` |
| Healthcheck port | `8000` |
| Healthcheck path | `/health` |
| Start Command | On |
| CMD | `--model Qwen/Qwen2.5-1.5B-Instruct --gpu-memory-utilization 0.9 --model-loader-extra-config '{"enable_multithread_load": true}'` |
| Environment variables | `HF_TOKEN`, set to your Hugging Face User Access Token |
| Scaling | Minimum number of replicas `1`, so each deployment stays ready during the test |

Wait until both deployments are healthy, then copy each one's **Endpoint Address** from its
deployment page. You paste them in Step 3.

For an explanation of each setting and a test request, see the
[vLLM quickstart](https://docs.verda.com/containers/tutorials/deploy-with-vllm-quick/). To serve the model with a different framework,
see the other [Serverless Container tutorials](https://docs.verda.com/containers/tutorials/), and keep the model the same on both
deployments.

***

## Step 2: Prepare the gateway instance

**Log in to the [Verda cloud console](https://console.verda.com/signin) and create a CPU
instance** with these settings:

| Setting | Value | Notes |
|---|---|---|
| Instance type | `CPU.4V.16G` | 4 vCPU and 16 GB RAM. |
| Image | Ubuntu 26.04 CUDA 13.1 + Docker | Docker and Docker Compose are preinstalled |
| SSH key | Your registered key | |

This size handled the short verification workload in this guide. Measure your own request
volume and response sizes before treating it as a production sizing recommendation.

Connect over SSH.

**Create the working directory** with a folder for test results and make your user its owner:

```bash
## New files are readable by you only
umask 077
## /opt belongs to root so creating the directory needs sudo
sudo mkdir -p /opt/llm-gateway/evidence
## Hand the directory to your user
sudo chown -R "$USER": /opt/llm-gateway
cd /opt/llm-gateway
```

This directory will hold the gateway credentials and test keys, which is why `umask 077`
keeps new files private. `evidence/` collects the responses and summaries saved by the tests.

**Check that the tools the guide uses are installed:**

```bash
docker --version
docker compose version
python3 --version
curl --version
```

Each command should print a version number. If Docker or Compose is missing, complete
[Docker's Ubuntu installation](https://docs.docker.com/engine/install/ubuntu/) before
continuing.

### Pull the gateway and database images

Download the LiteLLM and PostgreSQL 16 images:

```bash
sudo docker pull docker.litellm.ai/berriai/litellm:main-stable
sudo docker pull postgres:16
```

`main-stable` is LiteLLM's stable release channel, as used in its
[official quickstart Compose file](https://github.com/BerriAI/litellm/blob/main/docker/docker-compose.quickstart.yml).
Both tags move to newer builds over time, so the image you pull may be newer than the one this
guide was tested with. `postgres:16` keeps PostgreSQL on major version 16, which matters
because a database created by one major version cannot be opened by another.

***

## Step 3: Set the Endpoint Addresses and credentials

LiteLLM needs to know where each model endpoint is and which inference key to use. It also
needs a master key for administration, a salt key for encryption and a password for
PostgreSQL. Docker Compose reads those values from a private `.env` file.

The setup script below prompts for the two Endpoint Addresses and Inference API Keys. It then generates
the local credentials. It checks that the images were downloaded and refuses to replace an
existing `.env`, so running it again cannot silently change the secrets used by your
database.

**Create `configure.py`** in `/opt/llm-gateway` by copying the whole block below into the
terminal. `cat > ... <<'EOF'` writes every line up to the closing `EOF` into the file and the
quotes around `'EOF'` keep the shell from changing any `$` signs in the code. The other files
in this guide are created the same way.

```bash title="configure.py"
# configure.py
cat > /opt/llm-gateway/configure.py <<'EOF'
"""Create private Compose settings for a fresh gateway. Never overwrite secrets."""
import getpass
import os
import re
import secrets
import subprocess
from pathlib import Path
from urllib.parse import urlsplit

IMAGES = {
    "LITELLM_IMAGE": "docker.litellm.ai/berriai/litellm:main-stable",
    "POSTGRES_IMAGE": "postgres:16",
}


def read_base(label):
    value = input(f"Endpoint {label}: paste its Endpoint Address followed by /v1: ").strip().rstrip("/")
    parsed = urlsplit(value)
    if (
        parsed.scheme != "https" or not parsed.hostname
        or parsed.username or parsed.password or parsed.query or parsed.fragment
        or not parsed.path.endswith("/v1")
        or any(c.isspace() or c in "'\"\\$" for c in value)
    ):
        raise SystemExit("Use the HTTPS Endpoint Address followed by /v1, without credentials, query, or fragment.")
    return value


def read_key(label):
    value = getpass.getpass(f"Endpoint {label}: paste its Inference API Key (hidden): ").strip()
    if not re.fullmatch(r"[A-Za-z0-9._~+/\-]+=*", value):
        raise SystemExit("Expected a non-empty Bearer token; paste the token only, without 'Bearer '.")
    return value


def main():
    os.umask(0o077)
    if Path(".env").exists():
        raise SystemExit(".env already exists. Keep it; this script does not rotate existing secrets.")
    for reference in IMAGES.values():
        subprocess.run(["docker", "image", "inspect", reference], check=True, stdout=subprocess.DEVNULL)
    values = dict(IMAGES)
    values.update(
        POSTGRES_PASSWORD=secrets.token_hex(24),
        LITELLM_MASTER_KEY="sk-" + secrets.token_hex(32),
        LITELLM_SALT_KEY="sk-" + secrets.token_hex(32),
    )
    for label in ("A", "B"):
        values[f"VERDA_{label}_API_BASE"] = read_base(label)
        values[f"VERDA_{label}_KEY"] = read_key(label)
    if values["VERDA_A_API_BASE"] == values["VERDA_B_API_BASE"]:
        raise SystemExit("A and B must be different Endpoint Addresses.")
    # Validated endpoint values are single-quoted for literal Compose dotenv parsing.
    lines = [f"{k}='{v}'" if k.startswith("VERDA_") else f"{k}={v}" for k, v in values.items()]
    with Path(".env").open("x") as handle:
        handle.write("\n".join(lines) + "\n")
    print("Created .env with private permissions. No keys were printed.")


if __name__ == "__main__":
    main()
EOF
```

**Run it:**

```bash
cd /opt/llm-gateway
sudo python3 configure.py
```

!!! example
    **Example: what the terminal asks**

    ```
    Endpoint A: paste its Endpoint Address followed by /v1: <endpoint-address-a>/v1
    Endpoint A: paste its Inference API Key (hidden):
    Endpoint B: paste its Endpoint Address followed by /v1: <endpoint-address-b>/v1
    Endpoint B: paste its Inference API Key (hidden):
    Created .env with private permissions. No keys were printed.
    ```

    For security reasons, the key prompts stay blank as you paste.

```bash
ls -l .env
```

Check that `.env` was created before continuing. You should see one file with permissions
`-rw-------` owned by `root`. If you see `No such file or directory` the script stopped before writing it.
Scroll up to the message it printed, fix that problem, and run the script again.

**These credentials have different jobs:**

| Credential | Who uses it | Why it exists |
|---|---|---|
| `VERDA_A_KEY`, `VERDA_B_KEY` | LiteLLM | The Inference API Keys you pasted for endpoints A and B. LiteLLM sends them to authenticate to each deployment |
| `LITELLM_MASTER_KEY` | The operator | Create and manage gateway keys |
| Virtual key (created later) | An application | Call the allowed model through the gateway |
| `LITELLM_SALT_KEY` | LiteLLM | Encrypt stored secrets |
| `POSTGRES_PASSWORD` | LiteLLM and PostgreSQL | Authenticate the database connection |

**Applications receive virtual keys.** They do not need the Verda inference keys or the
master key. Keep `.env` private, preserve the salt for the lifetime of the database and
provision these values through your secret-management system for ongoing operation.

!!! warning
    Editing the password or salt in `.env` is not a complete rotation procedure. See [LiteLLM's database-backed setup](https://docs.litellm.ai/docs/proxy/docker_quick_start).

!!! info
    The gateway only needs **Inference API Keys**. Cloud API credentials used to manage Verda resources are not needed here. If a GPU container itself needs a secret such as a model-download token, set it on the deployment by following the [Serverless Container secrets guide](https://datacrunch-python.readthedocs.io/en/latest/examples/containers/secrets.html). Those secrets stay with the container and are not available on the gateway instance.

***

## Step 4: Define the model pool and services

We give both deployments the same client-facing model name. That tells LiteLLM they are two
places to serve the same request. Their deployment IDs remain different so that we can see
where traffic went.

**Create `config.yaml`** in `/opt/llm-gateway`:

```bash title="config.yaml"
# config.yaml
cat > /opt/llm-gateway/config.yaml <<'EOF'
model_list:
  - model_name: verda-qwen
    litellm_params:
      model: openai/Qwen/Qwen2.5-1.5B-Instruct
      api_base: os.environ/VERDA_A_API_BASE
      api_key: os.environ/VERDA_A_KEY
    model_info:
      id: verda-a
      mode: chat
      input_cost_per_token: 0
      output_cost_per_token: 0
      health_check_max_tokens: 8
      health_check_timeout: 10

  - model_name: verda-qwen
    litellm_params:
      model: openai/Qwen/Qwen2.5-1.5B-Instruct
      api_base: os.environ/VERDA_B_API_BASE
      api_key: os.environ/VERDA_B_KEY
    model_info:
      id: verda-b
      mode: chat
      input_cost_per_token: 0
      output_cost_per_token: 0
      health_check_max_tokens: 8
      health_check_timeout: 10

router_settings:
  routing_strategy: simple-shuffle
  num_retries: 2
  timeout: 30

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  database_url: os.environ/DATABASE_URL
  background_health_checks: true
  health_check_interval: 30
  enable_health_check_routing: true
EOF
```

`model_name` is the client-facing alias. The `openai/` prefix selects LiteLLM's
OpenAI-compatible adapter; the remainder must match the upstream model's served name. If you
chose another model, change both upstream model entries. See the
[configuration reference](https://docs.litellm.ai/docs/proxy/configs).

`simple-shuffle` picks one of the eligible deployments at random for each request. Because
the choice is random, a handful of requests can land unevenly on A and B; the split evens out
over many requests. The `model_info.id` values identify which deployment served a request. See
[load balancing](https://docs.litellm.ai/docs/proxy/load_balancing) and
[response headers](https://docs.litellm.ai/docs/proxy/response_headers).

The `input_cost_per_token: 0` and `output_cost_per_token: 0` lines deliberately make this a
**token-accounting example**: LiteLLM counts tokens but assigns them no price. They do not make
GPU use free, calculate your Verda bill, or enforce monetary budgets. LiteLLM documents that
setting both prices to zero bypasses its monetary budget enforcement. Configure an explicit
pricing policy if you need monetary accounting. See
[custom pricing](https://docs.litellm.ai/docs/proxy/custom_pricing).

### Define the containers

PostgreSQL makes virtual keys and usage records persistent. LiteLLM uses the database
connection supplied by Compose and reads the model pool from `config.yaml`.

Create `compose.yaml` in `/opt/llm-gateway`:

```bash title="compose.yaml"
# compose.yaml
cat > /opt/llm-gateway/compose.yaml <<'EOF'
name: llm-gateway

services:
  db:
    image: ${POSTGRES_IMAGE:?Missing POSTGRES_IMAGE}
    restart: unless-stopped
    environment:
      POSTGRES_USER: litellm
      POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?Missing database password}
      POSTGRES_DB: litellm
    volumes:
      - postgres_data:/var/lib/postgresql/data
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U litellm -d litellm"]
      interval: 5s
      timeout: 5s
      retries: 20

  litellm:
    image: ${LITELLM_IMAGE:?Missing LITELLM_IMAGE}
    restart: unless-stopped
    env_file:
      - .env
    environment:
      DATABASE_URL: postgresql://litellm:${POSTGRES_PASSWORD}@db:5432/litellm
    volumes:
      - ./config.yaml:/app/config.yaml:ro
    command:
      - "--config"
      - "/app/config.yaml"
      - "--host"
      - "0.0.0.0"
      - "--port"
      - "4000"
      - "--num_workers"
      - "1"
    ports:
      - "127.0.0.1:4000:4000"
    depends_on:
      db:
        condition: service_healthy
    healthcheck:
      test:
        - CMD
        - python
        - -c
        - "import urllib.request; urllib.request.urlopen('http://127.0.0.1:4000/health/readiness', timeout=5)"
      interval: 10s
      timeout: 10s
      retries: 18
      start_period: 60s

volumes:
  postgres_data:
EOF
```

The database uses a named Docker volume, so replacing its container does not discard its
data. That volume is local to this host; it is not a backup or a second copy on another
machine.

The gateway publishes port 4000 on `127.0.0.1`, so it is reachable only from the instance
itself during this walkthrough. PostgreSQL has no published host port. The single worker keeps
the first rate-limit test straightforward; shared state becomes necessary when adding workers
or replicas.

***

## Step 5: Start the gateway and check model health

**Start the two services** from `/opt/llm-gateway`. Compose waits for PostgreSQL to be healthy
before starting LiteLLM:

```bash
sudo docker compose config --quiet
sudo docker compose up -d --wait --wait-timeout 240
sudo docker compose ps
curl --silent --show-error --fail-with-body \
  --max-time 15 http://127.0.0.1:4000/health/readiness
```

Both services should become healthy and readiness should return HTTP 200. `config --quiet`
validates the files without printing resolved secrets.

Readiness checks the gateway and its database; it does not establish that either model
endpoint is healthy. The authenticated `/health` route reports model health. With background
checks enabled, it returns the latest cached probe results. See
[LiteLLM health checks](https://docs.litellm.ai/docs/proxy/health).

**Load the master key** for the administrative commands that follow, without printing it:

```bash
LITELLM_MASTER_KEY=$(sudo sed -n 's/^LITELLM_MASTER_KEY=//p' .env)
: "${LITELLM_MASTER_KEY:?Master key is missing}"
```

**Define this helper** in the same Bash session:

```bash
check_gateway_health() {
  local label="$1"
  curl --silent --show-error --fail-with-body --max-time 60 \
    http://127.0.0.1:4000/health \
    --header "Authorization: Bearer ${LITELLM_MASTER_KEY:?Master key missing}" \
    --output "evidence/health-${label}.json" || return

  python3 - "evidence/health-${label}.json" <<'PY'
import json
import sys
with open(sys.argv[1]) as handle:
    data = json.load(handle)
for name in ("healthy_endpoints", "unhealthy_endpoints"):
    endpoints = data.get(name, [])
    print(f"{name}: {len(endpoints)}")
    for endpoint in endpoints:
        print(" ", endpoint.get("api_base", "(URL not reported)"))
PY
}

check_gateway_health baseline
```

Wait for **two healthy endpoints and zero unhealthy endpoints** before continuing. If the
first background cycle has not completed, wait 30 seconds and call the helper again. A 200
status from `/health` alone is not enough; inspect its endpoint lists.

!!! info "If you reconnect"
    `LITELLM_MASTER_KEY`, `CLIENT_KEY` (from Step 7) and `check_gateway_health` exist only in the current terminal session. After logging out or opening a new terminal, commands fail with HTTP 401 *Malformed API Key*. Reload the keys from their files:

    ```bash
    cd /opt/llm-gateway
    LITELLM_MASTER_KEY=$(sudo sed -n 's/^LITELLM_MASTER_KEY=//p' .env)
    [ -f client-key.json ] && CLIENT_KEY=$(python3 -c 'import json; print(json.load(open("client-key.json"))["key"])')
    ```

    Then paste the `check_gateway_health` block above again before running Step 10.

***

## Step 6: Send the first request

We can now send a request through the gateway. The payload uses `verda-qwen`, so LiteLLM
chooses the underlying deployment.

**Create `gateway-request.json`** in `/opt/llm-gateway`:

```bash title="gateway-request.json"
# gateway-request.json
cat > /opt/llm-gateway/gateway-request.json <<'EOF'
{
  "model": "verda-qwen",
  "messages": [{"role": "user", "content": "Reply with one short greeting."}],
  "max_tokens": 32,
  "stream": false
}
EOF
```

**Send a request.** For this first check, use the master key already loaded in the shell. We create an
application key in the next step:

```bash
curl --silent --show-error --fail-with-body \
  --connect-timeout 5 --max-time 120 \
  http://127.0.0.1:4000/v1/chat/completions \
  --header "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --header "Content-Type: application/json" \
  --data-binary @gateway-request.json \
  --dump-header evidence/gateway-response.headers \
  --output evidence/gateway-response.json \
  --write-out 'Gateway request: HTTP %{http_code}\n'

python3 -m json.tool evidence/gateway-response.json
```

Expect HTTP 200, an assistant message, and a `usage` object containing input and output token
counts.

**Check that unauthenticated requests are rejected:**

```bash
curl --silent --show-error --max-time 30 \
  http://127.0.0.1:4000/v1/chat/completions \
  --header "Content-Type: application/json" \
  --data-binary @gateway-request.json \
  --output evidence/auth-missing-key.json \
  --write-out 'No API key: HTTP %{http_code}\n'

curl --silent --show-error --max-time 30 \
  http://127.0.0.1:4000/v1/chat/completions \
  --header "Authorization: Bearer sk-deliberately-invalid" \
  --header "Content-Type: application/json" \
  --data-binary @gateway-request.json \
  --output evidence/auth-invalid-key.json \
  --write-out 'Invalid API key: HTTP %{http_code}\n'
```

Both requests should return **HTTP 401**.

***

## Step 7: Give an application its own key

A virtual key lets you control one application without sharing an upstream inference
credential. Here we allow only `verda-qwen` and set a deliberately low limit of three requests
per minute, so we can see the rejection happen.

**Create the test key once.** It expires after 24 hours, and its returned secret is saved in
`client-key.json`. Running key generation again creates another key:

```bash
curl --silent --show-error --fail-with-body --max-time 30 \
  http://127.0.0.1:4000/key/generate \
  --header "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --header "Content-Type: application/json" \
  --data '{"key_alias":"llm-gateway-rpm-test","models":["verda-qwen"],"rpm_limit":3,"duration":"24h"}' \
  --output client-key.json \
  --write-out 'Create client key: HTTP %{http_code}\n'

CLIENT_KEY=$(python3 -c 'import json; print(json.load(open("client-key.json"))["key"])')
```

Use virtual keys in applications. Administrative keys can bypass limits, so using the master
key would not test the client policy. See
[virtual keys](https://docs.litellm.ai/docs/proxy/virtual_keys) and
[rate limits](https://docs.litellm.ai/docs/proxy/users#set-rate-limits).

**Send six requests** immediately with the new client key:

```bash
for i in 1 2 3 4 5 6; do
  curl --silent --show-error --connect-timeout 5 --max-time 30 \
    http://127.0.0.1:4000/v1/chat/completions \
    --header "Authorization: Bearer ${CLIENT_KEY:?Client key missing}" \
    --header "Content-Type: application/json" \
    --data-binary @gateway-request.json \
    --dump-header "evidence/rpm-${i}.headers" \
    --output "evidence/rpm-${i}.json" \
    --write-out "Request ${i}: HTTP %{http_code}, duration %{time_total}s\n"
done | tee evidence/rpm-summary.txt
```

Expect three HTTP 200 responses followed by HTTP 429 responses. The gateway accepted the
first three calls and rejected the rest at the configured limit. Run the loop only once. If you
need to repeat it, for example because the requests spanned a one-minute window boundary,
first create a new key with a new `key_alias` such as `llm-gateway-rpm-test-2`, and use that
alias in Step 8's query.

The reference stack uses one worker so its in-memory counters have one owner, and a restart
resets that state. Add shared Redis coordination before adding workers or gateway replicas;
PostgreSQL alone does not synchronize the rate limiter. See
[LiteLLM Redis requirements](https://docs.litellm.ai/docs/proxy/redis_requirements).

***

## Step 8: Check the tokens used

Each successful response includes `usage`: tokens in the prompt, tokens in the completion,
and their total. LiteLLM also records that usage in PostgreSQL, where it is associated with
the application key.

We check the count from both sides: first what the client received, then what the gateway
stored.

**What the client received.** Step 7 saved each response as `evidence/rpm-1.json` to
`evidence/rpm-6.json`. The following script opens those six files, skips the rejected ones,
and adds up the `usage` of the rest:

```bash
python3 - <<'PY'
import json
from pathlib import Path
names = ("prompt_tokens", "completion_tokens", "total_tokens")
totals = dict.fromkeys(names, 0)
successful = 0
for i in range(1, 7):
    data = json.loads(Path(f"evidence/rpm-{i}.json").read_text())
    if data.get("choices") and data.get("usage"):
        successful += 1
        for name in names:
            totals[name] += data["usage"].get(name, 0)
print("Responses with usage:", successful)
print(json.dumps(totals, indent=2))
PY
```

`range(1, 7)` means files 1 to 6, one per request sent in Step 7. If you sent a different
number of requests, change `7` to that number plus one.

In the reference test, it printed:

```
Responses with usage: 3
{
  "prompt_tokens": 105,
  "completion_tokens": 30,
  "total_tokens": 135
}
```

**What the gateway stored.** Now query PostgreSQL for the usage LiteLLM recorded against this
test key. Usage writes can be asynchronous, so give them a few seconds to finish first. The
query matches requests to the key by its alias, without displaying the key itself:

```bash
sudo docker compose exec -T db psql -U litellm -d litellm -v ON_ERROR_STOP=1 <<'SQL'
SELECT
  COUNT(*) FILTER (WHERE s.total_tokens > 0) AS requests_with_usage,
  COALESCE(SUM(s.prompt_tokens), 0) AS prompt_tokens,
  COALESCE(SUM(s.completion_tokens), 0) AS completion_tokens,
  COALESCE(SUM(s.total_tokens), 0) AS total_tokens
FROM "LiteLLM_SpendLogs" AS s
JOIN "LiteLLM_VerificationToken" AS k ON s.api_key = k.token
WHERE k.key_alias = 'llm-gateway-rpm-test';
SQL
```

In the reference test, it returned:

```
 requests_with_usage | prompt_tokens | completion_tokens | total_tokens
---------------------+---------------+-------------------+--------------
                   3 |           105 |                30 |          135
(1 row)
```

**The two totals match**: three successful requests, 105 prompt tokens plus 30 completion
tokens, 135 in total. The three rejected requests are stored as zero-token failure records, so
they do not count towards `requests_with_usage`. Your numbers depend on the model, prompt and
output, so compare your own two totals rather than expecting 135.

If you ran the Step 7 loop more than once with the same key, the database total includes every
run while the client total covers only the last one, so the two won't match. This read-only query matches the tested schema; recheck it when
upgrading LiteLLM.

Token counts describe model usage. They are separate from time-based infrastructure charges.
For logging configuration and retention, see
[LiteLLM spend tracking](https://docs.litellm.ai/docs/proxy/cost_tracking). This walkthrough
verifies non-streaming accounting; validate streaming separately if your application uses it.

***

## Step 9: Send traffic to both deployments

One successful response shows that the gateway can reach a model. To see whether both
endpoints serve traffic, we send 100 short requests with ten concurrent clients and count the
deployment IDs in the response headers.

**Create `load-test.py`** in `/opt/llm-gateway`. It reads the same `gateway-request.json` and
writes its results to `evidence/`:

```bash title="load-test.py"
# load-test.py
cat > /opt/llm-gateway/load-test.py <<'EOF'
import json
import math
import statistics
import sys
import time
from collections import Counter
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
from urllib.error import HTTPError
from urllib.request import Request, urlopen

label = sys.argv[1] if len(sys.argv) > 1 else "baseline"
count = int(sys.argv[2]) if len(sys.argv) > 2 else 100
workers = int(sys.argv[3]) if len(sys.argv) > 3 else 10

key = json.loads(Path("load-key.json").read_text())["key"]
payload = Path("gateway-request.json").read_bytes()
stamp = time.strftime("%Y%m%dT%H%M%SZ", time.gmtime())
folder = Path("evidence") / f"{label}-{stamp}"
folder.mkdir(parents=True, exist_ok=True)

def send_request(number):
    started = time.perf_counter()
    row = {"number": number, "status": 0}
    request = Request(
        "http://127.0.0.1:4000/v1/chat/completions",
        data=payload,
        headers={
            "Authorization": f"Bearer {key}",
            "Content-Type": "application/json",
        },
    )
    try:
        with urlopen(request, timeout=120) as response:
            row["status"] = response.status
            row["endpoint"] = response.headers.get("x-litellm-model-id")
            row["call_id"] = response.headers.get("x-litellm-call-id")
            row["retries"] = response.headers.get("x-litellm-attempted-retries")
            body = json.loads(response.read())
            row["response_id"] = body.get("id")
            row["usage"] = body.get("usage", {})
            if not body.get("choices") or not row["usage"]:
                row["error"] = "Missing completion or usage"
    except HTTPError as error:
        row["status"] = error.code
        row["error"] = "HTTPError"
    except Exception as error:
        row["error"] = type(error).__name__
    row["seconds"] = round(time.perf_counter() - started, 6)
    return row

started = time.perf_counter()
with ThreadPoolExecutor(max_workers=workers) as pool:
    rows = list(pool.map(send_request, range(1, count + 1)))
elapsed = time.perf_counter() - started

good = [r for r in rows if r["status"] == 200 and "error" not in r]
latencies = sorted(r["seconds"] for r in good)
summary = {
    "label": label,
    "requests": count,
    "concurrency": workers,
    "successful": len(good),
    "failed": count - len(good),
    "http_statuses": dict(Counter(str(r["status"]) for r in rows)),
    "successful_requests_by_endpoint": dict(
        Counter(r.get("endpoint") or "unknown" for r in good)
    ),
    "elapsed_seconds": round(elapsed, 3),
    "successful_requests_per_second": round(len(good) / elapsed, 2),
    "p50_seconds": round(statistics.median(latencies), 4) if good else None,
    "p95_seconds": round(latencies[math.ceil(0.95 * len(latencies)) - 1], 4) if good else None,
    "tokens": {
        name: sum(r["usage"].get(name, 0) for r in good)
        for name in ("prompt_tokens", "completion_tokens", "total_tokens")
    },
}
(folder / "requests.json").write_text(json.dumps(rows, indent=2))
(folder / "summary.json").write_text(json.dumps(summary, indent=2))
print(json.dumps(summary, indent=2))
print(f"Evidence saved in: {folder}")
EOF
```

**Run the test.** The script reads `load-key.json`. Create that file by issuing a separate test key with enough
allowance for this run, then start the test:

```bash
curl --silent --show-error --fail-with-body --max-time 30 \
  http://127.0.0.1:4000/key/generate \
  --header "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --header "Content-Type: application/json" \
  --data '{"key_alias":"llm-gateway-load-test","models":["verda-qwen"],"rpm_limit":1000,"tpm_limit":100000,"duration":"24h"}' \
  --output load-key.json

python3 load-test.py baseline 100 10
```

The printed summary reports success counts, latency, token totals, and the number of requests
served by each deployment. The timestamped directory holds both the summary and the
individual results. Check for:

- `successful: 100` and `failed: 0`.
- Both `verda-a` and `verda-b` in `successful_requests_by_endpoint`, adding up to 100.
- No `unknown` deployment IDs, which would mean the routing evidence is incomplete.

The split will not be exactly 50/50 and changes from run to run; what matters is that both
deployments receive requests. Latency depends on your model, GPU and prompt length, so use
`p50_seconds` and `p95_seconds` as your own baseline for later comparison. This short test
shows that traffic is distributed; it is not a throughput benchmark.

The load key also shows how to configure `tpm_limit`, a tokens-per-minute limit. Rejection at
that threshold was not part of the reference test.

***

## Step 10: Stop one endpoint and verify recovery

The next check answers a different question: can the same client-facing model keep serving
after one endpoint is taken out of service?

**First, distinguish the three health checks.** A healthy gateway process and a healthy model
endpoint are different things:

| Check | What it establishes |
|---|---|
| Verda container readiness | A replica is ready for work within its deployment |
| LiteLLM `/health/readiness` | The gateway and its database are ready |
| LiteLLM background model checks | Each configured model endpoint can complete a small inference request |

Here, the model probes run every 30 seconds with a ten-second timeout and an eight-token
output cap. Health-based routing removes a failing deployment from eligibility and admits it
again after a successful check. Both `background_health_checks` and
`enable_health_check_routing` must be enabled. See
[health-check-driven routing](https://docs.litellm.ai/docs/proxy/health_check_routing).

### Take A out of service

1. Confirm both endpoints are healthy and keep deployment B running.
2. In the Verda console, pause deployment A.
3. Run `check_gateway_health a-stopped`. Repeat after another health cycle until A is
   unhealthy and B is healthy.
4. Run the same load test:

    ```bash
    python3 load-test.py a-stopped 100 10
    ```

After health detection, expect 100 successful requests served by `verda-b`, with no successful
requests assigned to `verda-a`.

### Restore A

1. Resume deployment A in the Verda console and wait for its replica to be ready.
2. Run `check_gateway_health a-restored` until both endpoints are healthy again.
3. Run:

    ```bash
    python3 load-test.py recovered 100 10
    ```

Expect successful traffic on both deployment IDs again. Keep all three summaries, baseline,
A stopped and recovered, for comparison.

**This sequence verifies routing after the gateway has detected an outage.** It does not
prove that requests already in flight survive, nor that the interval between failure and
detection is error-free. Retries can extend latency and repeat upstream computation, so test
that transition under your application's workload.

If every deployment is unhealthy, LiteLLM documents that it can bypass the health filter and
keep attempting requests. A successful `/health` HTTP status therefore does not guarantee
usable model capacity. Monitor the endpoint counts and real completion failures.

!!! info
    The model probes send real inference traffic, so they may keep a deployment from scaling to zero. Check this against your scaling settings. For a latency-sensitive standby, keep capacity ready and account for its running cost.

Congratulations! You now have one model alias, per-application access control, persistent
token records and a way to verify which endpoint handled each request.

***

## Conclusion

You now have a LiteLLM gateway on a CPU instance that authenticates applications with their
own keys, enforces request limits, records token usage in PostgreSQL and spreads traffic
across two Serverless Container deployments with health-based failover.

**To clean up.** The test keys expire after 24 hours. Use the
[virtual-key management API](https://docs.litellm.ai/docs/proxy/virtual_keys) to revoke them
earlier, and keep any saved key files private. To stop the local services while keeping the
database volume:

```bash
cd /opt/llm-gateway
sudo docker compose stop
```

Stopping these services does not stop the instance or the two GPU deployments, which keep
billing while they exist. When you are finished,
[discontinue](https://docs.verda.com/cpu-and-gpu-instances/shutdown-hibernate-and-delete/) the instance and
delete both container deployments. Keep the database and its secrets if you intend to resume
the gateway.

***

## Related guides

- [Deploy with vLLM](https://docs.verda.com/containers/tutorials/deploy-with-vllm-quick/): create the model endpoints used by the
  gateway.
- [Serverless Container tutorials](https://docs.verda.com/containers/tutorials/): other inference deployment examples.
- [Scaling and health checks](https://docs.verda.com/containers/scaling/): replica readiness, queueing and scaling
  behaviour.
- [API credentials](https://docs.verda.com/welcome-to-verda/api-credentials/): create the Inference API Keys
  for the upstream endpoints.
- [Serverless Container secrets](https://datacrunch-python.readthedocs.io/en/latest/examples/containers/secrets.html):
  manage secrets used by GPU containers.
- [LiteLLM production deployment](https://docs.litellm.ai/docs/proxy/deploy): expand the
  reference gateway for your availability requirements.
