---
description: "Build a RAG service on Verda: retrieve passages from your own documents with pgvector and answer questions with an LLM on a Serverless Container, with a citation on every claim."
revision_date: 24.09.2026
---

# Tutorial: Retrieval-Augmented Generation (RAG) with pgvector

## What you'll build

In this guide we will build a retrieval-augmented generation (RAG) service: a language model
that answers questions using only passages retrieved from your own documents, with a
**citation** on every claim.

- You ask **"Am I charged for a volume when the instance is shut down?"**
- You get a written answer with **[1]**, **[2]** after each claim and a Sources list naming
  the page and section each one came from, so a reader can check it.

Where the Vector DB guide returned the matching rows, this guide reads them and answers.

Any self-hosted vector database works for this: **Qdrant**, **Weaviate**, **Milvus** or
**pgvector**. **This guide uses pgvector**, and the database and embedding deployment are the
ones from the [Vector DB guide](https://docs.verda.com/containers/tutorials/guide-vector-database/).

***

## Architecture

- Your documents live in **Object Storage** and are uploaded with the Verda CLI.
- Passages are turned into vectors by an embedding model on a **Serverless Container**.
- Vectors are stored and searched in **Postgres + pgvector** on a CPU instance with a
  Block Volume.
- Answers come from an LLM on a second **Serverless Container**, served by vLLM.

Two flows connect them. They share the database and the embedding deployment, but run at
different times.

**Ingestion** runs whenever your documents change:

<pre style="overflow-x: auto; white-space: pre;">

  +--------------------+
  |   Object Storage   |  your documents, as Markdown
  +---------+----------+
            | 1. read
            v
  +--------------------+   2. embed    +----------------------+
  |     ingest.py      | ------------> |   embedding model    |
  |   chunk > embed    | <------------ |                      |
  +---------+----------+    vectors    +----------------------+
            | 3. store                   GPU, Serverless Container
            v
  +----------------------------------+
  |      Postgres + pgvector         |  CPU instance + Block Volume
  +----------------------------------+

</pre>

**Query** runs once per question:

<pre style="overflow-x: auto; white-space: pre;">

        question
            |
            v
  +--------------------+   1. embed    +----------------------+
  |                    | ------------> |   embedding model    |
  |                    | <------------ |                      |
  |                    |    vector     +----------------------+
  |                    |
  |      query.py      |   2. search   +----------------------+
  |                    | ------------> | Postgres + pgvector  |
  |                    | <------------ |                      |
  |                    |  top-k chunks +----------------------+
  |                    |
  |                    |   3. prompt   +----------------------+
  |                    | ------------> |     LLM (vLLM)       |
  |                    | <------------ |                      |
  +---------+----------+    answer     +----------------------+
            |                            GPU, Serverless Container
            v
    answer + citations

</pre>

`ingest.py` and `query.py` are the two scripts you write, and both run on the CPU instance.
The LLM is only called at query time.

<div style="display: block !important; width: 100% !important; margin: 0 !important;">
<table style="display: table !important; width: 100% !important; margin: 0 !important; table-layout: fixed;">
<thead>
<tr><th>Component</th><th>Tier</th><th>Runs when</th></tr>
</thead>
<tbody>
<tr><td>Database</td><td>CPU.4V.16G instance</td><td>Continuously, while the instance is on</td></tr>
<tr><td>Data volume</td><td>NVMe Block Volume</td><td>Always, whether or not the instance runs</td></tr>
<tr><td>Embeddings</td><td>Serverless Container, L40S</td><td>On demand, scales to zero when idle</td></tr>
<tr><td>LLM</td><td>Serverless Container, L40S</td><td>On demand, scales to zero when idle</td></tr>
<tr><td>Documents</td><td>Object Storage</td><td>Always, billed on stored data</td></tr>
</tbody>
</table>
</div>

**Ingestion and querying must use the same embedding deployment.** Vectors from different
models cannot be compared, and nothing reports the mismatch. Both scripts import one embedding
function for that reason.

***

## Prerequisites

- Complete Steps 1, 2 and 4 of the [Vector DB guide](https://docs.verda.com/containers/tutorials/guide-vector-database/), so you have
  the CPU instance with its Block Volume, Postgres with pgvector 0.8 or newer, and the
  `embedder` deployment running. Its schema (Step 3) is not needed; this guide creates its own.
- Create an Object Storage access key.
- Have ready your [Inference API
  Key](../../welcome-to-verda/api-credentials.md#create-inference-api-keys) and Hugging Face
  token from the Vector DB guide.
- Install the [Verda CLI](https://docs.verda.com/cli/getting-started/) on the CPU instance from the Vector DB
  guide, and authenticate it with your Cloud API Client ID and Secret. Test it with
  `verda locations`; if it lists datacenters, the install and login are working.

Everything here is done in the Verda console, with the Verda CLI and over SSH.

No documents are needed to start: Step 1 creates a small sample corpus. Swap in your own
Markdown files later; the pipeline does not change.

***

## Step 1: Upload your documents

SSH into the CPU instance from the Vector DB guide.

### Create the sample corpus

The corpus is a folder of Markdown files. To have something to index, create three short pages
about Verda storage, billing and containers. In your home directory, create the folder:

```bash
mkdir -p corpus
```

Then create the three files. Each `cat` command writes one file; paste it whole, including
the closing `EOF` line:

```bash title="corpus/block-volumes.md"
cat > corpus/block-volumes.md <<'EOF'
# Block Volumes

A Block Volume is NVMe storage that exists independently of any instance, like an external
drive. You attach it to an instance to use it, detach it, and attach it to another instance.

## Attaching and resizing

A volume can be added when an instance is created or attached to an existing instance. It can
be expanded later without interrupting the instance. Volumes cannot be shrunk.

## Cloning a volume

A volume can be cloned within a datacenter or to another one, for backups or to seed a new
instance. The volume must be detached, or its instance shut down, while the clone is made.

## Deleting a volume

A deleted volume goes to the trash for 96 hours and can be restored during that time. To free
the storage quota and stop charges, permanently delete it from the Deleted volumes section.
Permanent deletion cannot be undone.
EOF
```

```bash title="corpus/pricing-and-billing.md"
cat > corpus/pricing-and-billing.md <<'EOF'
# Pricing and billing

Verda bills Pay As You Go in pre-paid 10-minute increments. If a resource is terminated before
the billed period ends, the unused portion is refunded in the next billing period.

## Storage charges

Storage is billed for as long as it exists, per GiB per month. A Block Volume is charged whether
or not it is attached to an instance, and whether or not that instance is running. Charges stop
only when the volume is permanently deleted. Restoring a volume from the trash costs the Pay As
You Go price for the time it was deleted.

## Shutting down and deleting instances

Shutting down an instance pauses it so that volumes can be attached or detached. A shut-down
instance continues to charge your account. Deleting an instance removes it together with the
volumes you choose to delete; any volume you keep continues to be billed and can be attached to
a new instance.
EOF
```

```bash title="corpus/serverless-containers.md"
cat > corpus/serverless-containers.md <<'EOF'
# Serverless Containers

Serverless Containers run any container image on GPU or CPU tiers and are billed only while
replicas are running.

## Scaling

Autoscaling applies whenever the maximum number of replicas is set higher than the minimum. By
default it scales on queue load, the queue length divided by the number of replicas. Scale-up
and scale-down delays control how long a threshold must hold before replicas change. With a
minimum of zero replicas, a deployment scales to zero when idle and costs nothing until the next
request, which then waits for a cold start.

## Health checks

Traffic is routed only to replicas that report healthy: an HTTP 200 from the healthcheck path.
During scale-down a replica receives SIGTERM and has 30 seconds to finish before SIGKILL.
EOF
```

Check that all three files are there:

```bash
ls corpus/
```

### Configure the CLI for Object Storage

Verda Object Storage is S3-compatible. We will manage it with the
[Verda CLI command](https://docs.verda.com/cli/object-storage/) `verda object-storage`, or its short form
`verda oss`. Object storage uses its own key, separate from the Cloud API credentials the CLI
already has. Create one and add it to the same profile with the following command:

```bash
verda oss configure
```

Name the profile `default`, then enter the access key and the secret. Accept the endpoint
(`https://objects.fin-03.verda.storage`) and region (`us-east-1`) with Enter.

Check that the key works:

```bash
verda oss ls
```

An empty list is the expected result, since you have no buckets yet. `AUTH_ERROR` means the
key was not accepted or is not in the active profile; `verda auth show` tells you which
profile that is.

### Create the bucket and upload

Bucket names are shared by everyone using the storage service, so pick one that is yours:
`rag-corpus-` followed by your name or team, in lowercase. Create the bucket:

```bash
verda oss mb s3://rag-corpus-<name>
```

You can also create it in the console instead; the upload below still uses the CLI.

Sync the folder into the bucket under a `raw/` prefix, then list it:

```bash
verda oss sync ./corpus/ s3://rag-corpus-<name>/raw/
verda oss ls s3://rag-corpus-<name> --recursive
```

The listing shows the three files under `raw/`:

```
✓ Listing objects...
  3 object(s) found

  LAST MODIFIED          SIZE          KEY
  -------------          ----          ---
  YYYY-MM-DD HH:MM:SS    852           raw/block-volumes.md
  YYYY-MM-DD HH:MM:SS    893           raw/pricing-and-billing.md
  YYYY-MM-DD HH:MM:SS    778           raw/serverless-containers.md
```

Re-running `sync` after you edit a document uploads only what changed. To use your own
documents, put the Markdown files in `corpus/` instead.

## Step 2: Deploy the LLM

The answers come from a chat model on a second Serverless Container. In the console, create a
deployment with these settings, which follow the [vLLM
tutorial](deploy-with-vllm-quick.md):

| Field | Value |
|---|---|
| Deployment name | `rag-llm` |
| GPU type | L40S 48GB, or RTX PRO 6000 if it has no capacity |
| Container image | `docker.io/vllm/vllm-openai:v0.27.1` |
| Exposed HTTP port and healthcheck port | `8000`, the port vLLM listens on |
| Healthcheck path | `/health`, the path vLLM exposes |
| Start command | Toggle on, which reveals the Entrypoint and CMD fields |
| Entrypoint | Leave empty, so the image's own entrypoint is used |
| CMD | `Qwen/Qwen3-8B --served-model-name rag-llm --max-model-len 16384` |
| Environment variables | Add one with the name `HF_TOKEN` and `<your-hugging-face-token>` as the value. Leave `HF_HOME` as the form sets it |

Any model that vLLM serves with a chat endpoint works here; the retrieval side does not depend
on it.

The first start takes longer than the embedder's, because the weights are about 16 GB. While
you wait, the system log shows the normal sequence: `Scheduled`, `Pulling`, `Pulled`,
`Created`, `Started`, then several `Startup probe failed: connection refused` entries while the
weights load. Those failures are expected, because nothing is listening on the port yet.

Once the deployment reports healthy, its page shows an **Endpoint Address**. Copy it, remove
the last `/`, and test it:

```bash
curl -sL -X POST <rag-llm-endpoint-address>/v1/chat/completions \
  --header "Authorization: Bearer <your-inference-api-key>" \
  --header 'Content-Type: application/json' \
  --data '{"model": "rag-llm", "messages": [{"role": "user", "content": "Say hello."}], "chat_template_kwargs": {"enable_thinking": false}}'
```

`enable_thinking: false` is there because Qwen3 is a reasoning model: without it, every reply
starts with a `<think>…</think>` section of the model's reasoning notes, before the actual
answer, which costs about a hundred extra tokens per reply. `query.py` in Step 5 sends its
questions to this endpoint, with the same setting.

A working response looks like this, trimmed:

```json
{"id":"chatcmpl-ba255b7f270fa7eb","object":"chat.completion","model":"rag-llm",
 "choices":[{"index":0,"message":{"role":"assistant","content":"Hello! How can I assist you today?"},
 "finish_reason":"stop"}],
 "usage":{"prompt_tokens":15,"total_tokens":27,"completion_tokens":12}}
```

`choices` is the list of answers; there is one unless you ask for more, so the text is at
`choices[0].message.content`, which is where `query.py` reads it.

## Step 3: Create the schema and roles

Back on the CPU instance. RAG needs two tables, one for the source files and one for their
chunks. Create a new database, `ragdb`, to hold them:

```bash
sudo -u postgres createdb ragdb
sudo -u postgres psql -d ragdb
```

```sql
CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE documents (
    doc_id       TEXT PRIMARY KEY,       -- object key without raw/ and .md
    title        TEXT NOT NULL,
    checksum     TEXT NOT NULL,          -- sha256 of the file, to skip unchanged docs
    ingested_at  TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE TABLE chunks (
    chunk_id     TEXT PRIMARY KEY,
    doc_id       TEXT NOT NULL REFERENCES documents(doc_id) ON DELETE CASCADE,
    chunk_index  INT  NOT NULL,
    heading      TEXT,                   -- the section the chunk came from
    content      TEXT NOT NULL,
    embedding    vector(1024) NOT NULL   -- Qwen3-Embedding-0.6B, same as the Vector DB guide
);
```

`documents` holds one row per source file with a checksum, so a re-run can skip files that have
not changed. `chunks` holds the passages that are searched, each pointing back to its document.

**Create two database roles**, so the ingester can write while the query side only reads. Pick
a password for each and substitute it for the placeholders, brackets included.

Still in psql:

```sql
CREATE ROLE rag_ingest LOGIN PASSWORD '<ingest-password>';
GRANT INSERT, UPDATE, DELETE, SELECT ON documents, chunks TO rag_ingest;

CREATE ROLE rag_query LOGIN PASSWORD '<query-password>';
GRANT SELECT ON documents, chunks TO rag_query;
```

Leave psql with `\q`.

### Set the environment variables

The scripts read everything they need from the environment. Use the same two passwords as
above; `HF_TOKEN` is the Hugging Face token; the two endpoints are the **Endpoint Address**
shown on the `embedder` and `rag-llm` deployment pages, with or without the trailing `/`.
If a password contains `@`, `:` or `/`, percent-encode it:

```bash
export VERDA_INFERENCE_KEY='<your-inference-api-key>'
export HF_TOKEN='<your-hugging-face-token>'
export EMBEDDER_ENDPOINT='<embedder-endpoint-address>'
export LLM_ENDPOINT='<rag-llm-endpoint-address>'
export VERDA_S3_ACCESS_KEY='<object-storage-access-key>'
export VERDA_S3_SECRET_KEY='<object-storage-secret-key>'
export RAG_BUCKET='rag-corpus-<name>'
export PG_INGEST_DSN='postgresql://rag_ingest:<ingest-password>@localhost/ragdb'
export PG_QUERY_DSN='postgresql://rag_query:<query-password>@localhost/ragdb'
```

These variables last only for the current shell. If you disconnect from the instance, run this
block again after you reconnect.

Check the database connection before going further:

```bash
psql "$PG_INGEST_DSN" -c "select 1"
```

The expected output is a one-row table containing `1`:

```
 ?column?
----------
        1
(1 row)
```

That means the `rag_ingest` role and password match. If you get `password authentication
failed`, the DSN password differs from the one you set; reset it in psql as `postgres` on
`ragdb` with `ALTER ROLE rag_ingest PASSWORD '<ingest-password>';` and run the check again.

## Step 4: Chunk and ingest

**Install the Python packages** in a virtual environment, in the same shell where you set the
variables:

```bash
sudo apt-get install -y python3-venv
python3 -m venv ~/vecenv
source ~/vecenv/bin/activate
pip install requests "psycopg[binary]" numpy boto3 tokenizers
```

### The ingester

The ingester splits each document into **chunks**, short passages that are embedded and
searched individually.

Save the following script as `ingest.py` in your home directory:

```python title="ingest.py"
# ingest.py
import hashlib, os, re, requests
import boto3, numpy as np, psycopg
from tokenizers import Tokenizer

BUCKET = os.environ["RAG_BUCKET"]
ENDPOINT = os.environ["EMBEDDER_ENDPOINT"].rstrip("/") + "/v1/embeddings"
KEY = "".join(os.environ["VERDA_INFERENCE_KEY"].split())   # strips stray whitespace
TARGET_TOKENS = 450        # chunk size; see "How chunking works"
BATCH = 64

tokenizer = Tokenizer.from_pretrained("Qwen/Qwen3-Embedding-0.6B")   # counts tokens like the model
FENCE = re.compile(r"^\s*(```|~~~)")
HEADING = re.compile(r"^#{1,6}\s+(.*)$")


def embed(texts: list[str]) -> np.ndarray:
    """The single source of truth for turning text into vectors. query.py imports it."""
    resp = requests.post(
        ENDPOINT,
        headers={"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"},
        json={"model": "embedder", "input": texts},
        timeout=300,               # generous: covers cold start
    )
    resp.raise_for_status()
    rows = [d["embedding"] for d in sorted(resp.json()["data"], key=lambda d: d["index"])]
    vecs = np.asarray(rows, dtype=np.float32)
    vecs /= np.linalg.norm(vecs, axis=1, keepdims=True)   # unit length, for cosine
    return vecs


def blocks(md: str):
    """Yield (heading, text) units: paragraphs, and fenced code blocks kept whole."""
    heading, buf, fence = "", [], None
    for line in md.split("\n") + [""]:
        if fence is None and FENCE.match(line):
            if buf:
                yield heading, "\n".join(buf); buf = []
            fence = FENCE.match(line).group(1); buf = [line]
        elif fence and line.strip().startswith(fence):
            buf.append(line); yield heading, "\n".join(buf); buf, fence = [], None
        elif fence:
            buf.append(line)
        elif HEADING.match(line):
            if buf:
                yield heading, "\n".join(buf); buf = []
            heading = HEADING.match(line).group(1).strip()
        elif line.strip() == "":
            if buf:
                yield heading, "\n".join(buf); buf = []
        else:
            buf.append(line)


def chunk(md: str):
    """Pack consecutive units under one heading into chunks of about TARGET_TOKENS."""
    chunks, cur, cur_heading, cur_tokens = [], [], None, 0
    for heading, text in blocks(md):
        n = len(tokenizer.encode(text).ids)
        if cur and (heading != cur_heading or cur_tokens + n > TARGET_TOKENS):
            chunks.append((cur_heading, "\n\n".join(cur))); cur, cur_tokens = [], 0
        cur.append(text); cur_heading = heading; cur_tokens += n
    if cur:
        chunks.append((cur_heading, "\n\n".join(cur)))
    return chunks


def ingest():
    s3 = boto3.client(
        "s3", endpoint_url="https://objects.fin-03.verda.storage", region_name="us-east-1",
        aws_access_key_id=os.environ["VERDA_S3_ACCESS_KEY"],
        aws_secret_access_key=os.environ["VERDA_S3_SECRET_KEY"],
    )
    conn = psycopg.connect(os.environ["PG_INGEST_DSN"])
    pages = s3.get_paginator("list_objects_v2").paginate(Bucket=BUCKET, Prefix="raw/")
    for obj in (o for p in pages for o in p.get("Contents", []) if o["Key"].endswith(".md")):
        key = obj["Key"]
        text = s3.get_object(Bucket=BUCKET, Key=key)["Body"].read().decode("utf-8")
        doc_id = key[len("raw/"):-len(".md")]
        checksum = hashlib.sha256(text.encode()).hexdigest()

        with conn.cursor() as cur:
            cur.execute("SELECT checksum FROM documents WHERE doc_id = %s", (doc_id,))
            if (row := cur.fetchone()) and row[0] == checksum:
                print(f"{doc_id}: unchanged"); continue

            title = next((m.group(1) for m in map(HEADING.match, text.split("\n")) if m), doc_id)
            cur.execute(
                """INSERT INTO documents (doc_id, title, checksum) VALUES (%s, %s, %s)
                   ON CONFLICT (doc_id) DO UPDATE
                   SET title = EXCLUDED.title, checksum = EXCLUDED.checksum, ingested_at = now()""",
                (doc_id, title, checksum),
            )
            cur.execute("DELETE FROM chunks WHERE doc_id = %s", (doc_id,))   # replace wholesale

            parts = chunk(text)
            for i in range(0, len(parts), BATCH):
                batch = parts[i:i + BATCH]
                vecs = embed([f"{h}\n\n{c}" if h else c for h, c in batch])
                cur.executemany(
                    """INSERT INTO chunks (chunk_id, doc_id, chunk_index, heading, content, embedding)
                       VALUES (%s, %s, %s, %s, %s, %s)""",
                    [(f"{doc_id}#{i + j}", doc_id, i + j, h, c, v.tolist())
                     for j, ((h, c), v) in enumerate(zip(batch, vecs))],
                )
            conn.commit()
            print(f"{doc_id}: {len(parts)} chunks")


if __name__ == "__main__":
    ingest()
```

Four functions do the work:

- `embed()` turns text into vectors. It is the function from the Vector DB guide.
- `blocks()` and `chunk()` split a Markdown file into chunks, one section at a time, keeping
  code blocks whole. [How chunking works](#how-chunking-works) explains the rules and how to
  tune the size.
- `ingest()` reads every `.md` file under `raw/` in the bucket, skips files that have not
  changed since the last run, and replaces the chunks of any file that has.

Run the script:

```bash
python3 ingest.py
```

The output shows one line per file:

```
block-volumes: 4 chunks
pricing-and-billing: 3 chunks
serverless-containers: 3 chunks
```

The numbers match the sections in each file: `block-volumes.md` has an intro and three
headings, so four chunks. The sections are short, so none needed splitting further.

### Build the index

Now that the rows are loaded, we will create the index. With this index, Postgres finds the
nearest vectors without comparing the question against every chunk. Building it after the load,
rather than before, means the inserts did not have to update it as they went.

Run this as the `postgres` superuser:

```bash
sudo -u postgres psql -d ragdb
```

```sql
SET maintenance_work_mem = '2GB';

CREATE INDEX chunks_embedding_idx
    ON chunks USING hnsw (embedding vector_cosine_ops)
    WITH (m = 16, ef_construction = 64);
```

Leave psql with `\q`.

## Step 5: Ask a question

Save this as `query.py` **in the same directory as `ingest.py`**, since it imports `embed()`
from it. Keep the venv active and the environment variables from Step 3 set:

```python title="query.py"
# query.py
import os, requests, psycopg
from ingest import embed

LLM = os.environ["LLM_ENDPOINT"].rstrip("/") + "/v1/chat/completions"
KEY = "".join(os.environ["VERDA_INFERENCE_KEY"].split())
conn = psycopg.connect(os.environ["PG_QUERY_DSN"])

SQL = """
SELECT c.heading, c.content, d.title, d.doc_id,
       1 - (c.embedding <=> %(q)s::vector) AS similarity
FROM chunks c JOIN documents d ON d.doc_id = c.doc_id
ORDER BY c.embedding <=> %(q)s::vector
LIMIT %(k)s
"""

SYSTEM = """Answer the question using only the numbered sources below.
Put the source number in square brackets after each claim, like [2].
If the sources do not contain the answer, say so. Do not invent source numbers."""


def retrieve(question: str, k: int = 6):
    qvec = embed([question])[0].tolist()
    with conn.cursor() as cur:
        cur.execute("SET LOCAL hnsw.ef_search = 100")
        cur.execute(SQL, {"q": qvec, "k": k})
        cols = [d.name for d in cur.description]
        return [dict(zip(cols, r)) for r in cur.fetchall()]


def answer(question: str):
    hits = retrieve(question)
    sources = "\n\n".join(f"[{i}] {h['title']} — {h['heading']}\n{h['content']}"
                          for i, h in enumerate(hits, 1))
    resp = requests.post(
        LLM,
        headers={"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"},
        json={"model": "rag-llm", "temperature": 0.2,
              "chat_template_kwargs": {"enable_thinking": False},   # no <think> preamble
              "messages": [{"role": "system", "content": SYSTEM},
                           {"role": "user", "content": f"Sources:\n\n{sources}\n\nQuestion: {question}"}]},
        timeout=300,
    )
    resp.raise_for_status()
    return resp.json()["choices"][0]["message"]["content"], hits


if __name__ == "__main__":
    import sys
    text, hits = answer(" ".join(sys.argv[1:]) or "Am I charged for a volume when the instance is shut down?")
    print(text, "\nSources")
    for i, h in enumerate(hits, 1):
        print(f"  [{i}] {h['title']} — {h['heading']}  ({h['doc_id']}.md, {h['similarity']:.2f})")
```

`retrieve()` is the search from the Vector DB guide, joined to `documents` for the title.
`answer()` numbers the retrieved chunks, tells the model to cite by number, and prints the
numbered list itself, so a citation always resolves to a real chunk.

Run the script with a question:

```bash
python3 query.py "Am I charged for a volume when the instance is shut down?"
```

On the sample corpus, the output is:

```
Yes, you are charged for a volume when the instance is shut down. A Block Volume is charged
whether or not it is attached to an instance, and whether or not that instance is running [2].
Sources
  [1] Pricing and billing — Shutting down and deleting instances  (pricing-and-billing.md, 0.75)
  [2] Pricing and billing — Storage charges  (pricing-and-billing.md, 0.72)
  [3] Pricing and billing — Pricing and billing  (pricing-and-billing.md, 0.61)
  [4] Block Volumes — Attaching and resizing  (block-volumes.md, 0.61)
  [5] Block Volumes — Deleting a volume  (block-volumes.md, 0.59)
  [6] Block Volumes — Cloning a volume  (block-volumes.md, 0.56)
```

The claim points at a section of one of the three files you created. All six retrieved chunks
are listed with their similarity to the question; the model used only the one it needed.

Now ask something the documents do not cover:

```bash
python3 query.py "Am I a cat?"
```

```
None of the provided sources address the question "Am I a cat?" Therefore, the answer cannot
be determined from the given sources.
Sources
  [1] Serverless Containers — Serverless Containers  (serverless-containers.md, 0.24)
  ...
```

The model says so instead of guessing, because the system prompt tells it to answer from the
sources only. The similarity scores, around 0.2 instead of 0.7, show that nothing relevant
was retrieved.

The score after each source runs from 0 to 1:

- **1** means the chunk has the same meaning as the question.
- **0** means the chunk is unrelated to it.

Congratulations! You now have a working RAG service.

***

## After it works

### When an answer is wrong

A citation tells you which source to check; it does not guarantee the model read that source
correctly. Look at the Sources list first. If the passage that answers the question is not
among the six, the model never saw it, and the fix is in chunking or retrieval rather than in
the prompt.

### How chunking works

The embedding model turns a piece of text into one vector that stands for its meaning as a
whole, and search returns the vectors closest to the question's. A vector for a whole page
captures what the page is about, not the details inside it; a vector for one section does
capture that section's details, so it matches a question about them. Documents
are therefore split into **chunks** of a few hundred tokens, each embedded on its own, and it
is chunks that are searched and handed to the LLM.

Two rules matter when your documents are Markdown files with headings and text blocks, as most
product docs are, and `ingest.py` applies both:

- **Split on headings**, so every chunk carries its section heading. The heading helps
  retrieval and is what the citation shows.
- **Never cut a command or code block in two.** It is added to a chunk whole or not at all,
  because half a command looks like a complete answer but fails when run.

The chunk size is `TARGET_TOKENS` at the top of `ingest.py`. 450 suits documentation pages:
smaller chunks are more precise but lose surrounding context, larger ones keep context but
match specific questions less well. Go lower for FAQ-style content, higher for long narrative
text, and re-run `ingest.py` on the whole corpus after any change.

### Updating the corpus

Edit or add Markdown files in `corpus/`, then run `verda oss sync` and `python3 ingest.py`
again. Files that have not changed (their checksum matches the one stored in `documents`) are
skipped, so only what changed is embedded.

***

## Conclusion

You now have documents in Object Storage, their passages indexed in pgvector, an embedding
model and an LLM on Serverless Containers, and a script that turns a question into a cited
answer.

**To clean up.** The instance, the volume and the bucket keep billing while they exist. When you
are finished, [discontinue](https://docs.verda.com/cpu-and-gpu-instances/shutdown-hibernate-and-delete/) the
instance, [permanently delete](https://docs.verda.com/storage/deleting-storage/) the volume, delete both
container deployments, and remove the bucket together with its files:

```bash
verda oss rb s3://rag-corpus-<name> --force
```

`--force` deletes everything in the bucket before removing it, and asks for confirmation first.

***

## Related guides

- [Vector DB guide](https://docs.verda.com/containers/tutorials/guide-vector-database/): the database and embedding deployment this
  guide builds on, with sizing, thresholds and backups.
- [Object Storage CLI](https://docs.verda.com/cli/object-storage/): the full command reference.
- [Serverless Containers storage](https://docs.verda.com/containers/storage/): scratch disk and shared filesystems, for
  caching model weights across cold starts.
