Skip to content

Tutorial: Retrieval-Augmented Generation (RAG) with pgvector

What you'll build

In this guide we will build a retrieval-augmented generation (RAG) service: a language model that answers questions using only passages retrieved from your own documents, with a citation on every claim.

  • You ask "Am I charged for a volume when the instance is shut down?"
  • You get a written answer with [1], [2] after each claim and a Sources list naming the page and section each one came from, so a reader can check it.

Where the Vector DB guide returned the matching rows, this guide reads them and answers.

Any self-hosted vector database works for this: Qdrant, Weaviate, Milvus or pgvector. This guide uses pgvector, and the database and embedding deployment are the ones from the Vector DB guide.


Architecture

  • Your documents live in Object Storage and are uploaded with the Verda CLI.
  • Passages are turned into vectors by an embedding model on a Serverless Container.
  • Vectors are stored and searched in Postgres + pgvector on a CPU instance with a Block Volume.
  • Answers come from an LLM on a second Serverless Container, served by vLLM.

Two flows connect them. They share the database and the embedding deployment, but run at different times.

Ingestion runs whenever your documents change:


  +--------------------+
  |   Object Storage   |  your documents, as Markdown
  +---------+----------+
            | 1. read
            v
  +--------------------+   2. embed    +----------------------+
  |     ingest.py      | ------------> |   embedding model    |
  |   chunk > embed    | <------------ |                      |
  +---------+----------+    vectors    +----------------------+
            | 3. store                   GPU, Serverless Container
            v
  +----------------------------------+
  |      Postgres + pgvector         |  CPU instance + Block Volume
  +----------------------------------+

Query runs once per question:


        question
            |
            v
  +--------------------+   1. embed    +----------------------+
  |                    | ------------> |   embedding model    |
  |                    | <------------ |                      |
  |                    |    vector     +----------------------+
  |                    |
  |      query.py      |   2. search   +----------------------+
  |                    | ------------> | Postgres + pgvector  |
  |                    | <------------ |                      |
  |                    |  top-k chunks +----------------------+
  |                    |
  |                    |   3. prompt   +----------------------+
  |                    | ------------> |     LLM (vLLM)       |
  |                    | <------------ |                      |
  +---------+----------+    answer     +----------------------+
            |                            GPU, Serverless Container
            v
    answer + citations

ingest.py and query.py are the two scripts you write, and both run on the CPU instance. The LLM is only called at query time.

ComponentTierRuns when
DatabaseCPU.4V.16G instanceContinuously, while the instance is on
Data volumeNVMe Block VolumeAlways, whether or not the instance runs
EmbeddingsServerless Container, L40SOn demand, scales to zero when idle
LLMServerless Container, L40SOn demand, scales to zero when idle
DocumentsObject StorageAlways, billed on stored data

Ingestion and querying must use the same embedding deployment. Vectors from different models cannot be compared, and nothing reports the mismatch. Both scripts import one embedding function for that reason.


Prerequisites

  • Complete Steps 1, 2 and 4 of the Vector DB guide, so you have the CPU instance with its Block Volume, Postgres with pgvector 0.8 or newer, and the embedder deployment running. Its schema (Step 3) is not needed; this guide creates its own.
  • Create an Object Storage access key.
  • Have ready your Inference API Key and Hugging Face token from the Vector DB guide.
  • Install the Verda CLI on the CPU instance from the Vector DB guide, and authenticate it with your Cloud API Client ID and Secret. Test it with verda locations; if it lists datacenters, the install and login are working.

Everything here is done in the Verda console, with the Verda CLI and over SSH.

No documents are needed to start: Step 1 creates a small sample corpus. Swap in your own Markdown files later; the pipeline does not change.


Step 1: Upload your documents

SSH into the CPU instance from the Vector DB guide.

Create the sample corpus

The corpus is a folder of Markdown files. To have something to index, create three short pages about Verda storage, billing and containers. In your home directory, create the folder:

mkdir -p corpus

Then create the three files. Each cat command writes one file; paste it whole, including the closing EOF line:

corpus/block-volumes.md
cat > corpus/block-volumes.md <<'EOF'
# Block Volumes

A Block Volume is NVMe storage that exists independently of any instance, like an external
drive. You attach it to an instance to use it, detach it, and attach it to another instance.

## Attaching and resizing

A volume can be added when an instance is created or attached to an existing instance. It can
be expanded later without interrupting the instance. Volumes cannot be shrunk.

## Cloning a volume

A volume can be cloned within a datacenter or to another one, for backups or to seed a new
instance. The volume must be detached, or its instance shut down, while the clone is made.

## Deleting a volume

A deleted volume goes to the trash for 96 hours and can be restored during that time. To free
the storage quota and stop charges, permanently delete it from the Deleted volumes section.
Permanent deletion cannot be undone.
EOF
corpus/pricing-and-billing.md
cat > corpus/pricing-and-billing.md <<'EOF'
# Pricing and billing

Verda bills Pay As You Go in pre-paid 10-minute increments. If a resource is terminated before
the billed period ends, the unused portion is refunded in the next billing period.

## Storage charges

Storage is billed for as long as it exists, per GiB per month. A Block Volume is charged whether
or not it is attached to an instance, and whether or not that instance is running. Charges stop
only when the volume is permanently deleted. Restoring a volume from the trash costs the Pay As
You Go price for the time it was deleted.

## Shutting down and deleting instances

Shutting down an instance pauses it so that volumes can be attached or detached. A shut-down
instance continues to charge your account. Deleting an instance removes it together with the
volumes you choose to delete; any volume you keep continues to be billed and can be attached to
a new instance.
EOF
corpus/serverless-containers.md
cat > corpus/serverless-containers.md <<'EOF'
# Serverless Containers

Serverless Containers run any container image on GPU or CPU tiers and are billed only while
replicas are running.

## Scaling

Autoscaling applies whenever the maximum number of replicas is set higher than the minimum. By
default it scales on queue load, the queue length divided by the number of replicas. Scale-up
and scale-down delays control how long a threshold must hold before replicas change. With a
minimum of zero replicas, a deployment scales to zero when idle and costs nothing until the next
request, which then waits for a cold start.

## Health checks

Traffic is routed only to replicas that report healthy: an HTTP 200 from the healthcheck path.
During scale-down a replica receives SIGTERM and has 30 seconds to finish before SIGKILL.
EOF

Check that all three files are there:

ls corpus/

Configure the CLI for Object Storage

Verda Object Storage is S3-compatible. We will manage it with the Verda CLI command verda object-storage, or its short form verda oss. Object storage uses its own key, separate from the Cloud API credentials the CLI already has. Create one and add it to the same profile with the following command:

verda oss configure

Name the profile default, then enter the access key and the secret. Accept the endpoint (https://objects.fin-03.verda.storage) and region (us-east-1) with Enter.

Check that the key works:

verda oss ls

An empty list is the expected result, since you have no buckets yet. AUTH_ERROR means the key was not accepted or is not in the active profile; verda auth show tells you which profile that is.

Create the bucket and upload

Bucket names are shared by everyone using the storage service, so pick one that is yours: rag-corpus- followed by your name or team, in lowercase. Create the bucket:

verda oss mb s3://rag-corpus-<name>

You can also create it in the console instead; the upload below still uses the CLI.

Sync the folder into the bucket under a raw/ prefix, then list it:

verda oss sync ./corpus/ s3://rag-corpus-<name>/raw/
verda oss ls s3://rag-corpus-<name> --recursive

The listing shows the three files under raw/:

✓ Listing objects...
  3 object(s) found

  LAST MODIFIED          SIZE          KEY
  -------------          ----          ---
  YYYY-MM-DD HH:MM:SS    852           raw/block-volumes.md
  YYYY-MM-DD HH:MM:SS    893           raw/pricing-and-billing.md
  YYYY-MM-DD HH:MM:SS    778           raw/serverless-containers.md

Re-running sync after you edit a document uploads only what changed. To use your own documents, put the Markdown files in corpus/ instead.

Step 2: Deploy the LLM

The answers come from a chat model on a second Serverless Container. In the console, create a deployment with these settings, which follow the vLLM tutorial:

Field Value
Deployment name rag-llm
GPU type L40S 48GB, or RTX PRO 6000 if it has no capacity
Container image docker.io/vllm/vllm-openai:v0.27.1
Exposed HTTP port and healthcheck port 8000, the port vLLM listens on
Healthcheck path /health, the path vLLM exposes
Start command Toggle on, which reveals the Entrypoint and CMD fields
Entrypoint Leave empty, so the image's own entrypoint is used
CMD Qwen/Qwen3-8B --served-model-name rag-llm --max-model-len 16384
Environment variables Add one with the name HF_TOKEN and <your-hugging-face-token> as the value. Leave HF_HOME as the form sets it

Any model that vLLM serves with a chat endpoint works here; the retrieval side does not depend on it.

The first start takes longer than the embedder's, because the weights are about 16 GB. While you wait, the system log shows the normal sequence: Scheduled, Pulling, Pulled, Created, Started, then several Startup probe failed: connection refused entries while the weights load. Those failures are expected, because nothing is listening on the port yet.

Once the deployment reports healthy, its page shows an Endpoint Address. Copy it, remove the last /, and test it:

curl -sL -X POST <rag-llm-endpoint-address>/v1/chat/completions \
  --header "Authorization: Bearer <your-inference-api-key>" \
  --header 'Content-Type: application/json' \
  --data '{"model": "rag-llm", "messages": [{"role": "user", "content": "Say hello."}], "chat_template_kwargs": {"enable_thinking": false}}'

enable_thinking: false is there because Qwen3 is a reasoning model: without it, every reply starts with a <think>…</think> section of the model's reasoning notes, before the actual answer, which costs about a hundred extra tokens per reply. query.py in Step 5 sends its questions to this endpoint, with the same setting.

A working response looks like this, trimmed:

{"id":"chatcmpl-ba255b7f270fa7eb","object":"chat.completion","model":"rag-llm",
 "choices":[{"index":0,"message":{"role":"assistant","content":"Hello! How can I assist you today?"},
 "finish_reason":"stop"}],
 "usage":{"prompt_tokens":15,"total_tokens":27,"completion_tokens":12}}

choices is the list of answers; there is one unless you ask for more, so the text is at choices[0].message.content, which is where query.py reads it.

Step 3: Create the schema and roles

Back on the CPU instance. RAG needs two tables, one for the source files and one for their chunks. Create a new database, ragdb, to hold them:

sudo -u postgres createdb ragdb
sudo -u postgres psql -d ragdb
CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE documents (
    doc_id       TEXT PRIMARY KEY,       -- object key without raw/ and .md
    title        TEXT NOT NULL,
    checksum     TEXT NOT NULL,          -- sha256 of the file, to skip unchanged docs
    ingested_at  TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE TABLE chunks (
    chunk_id     TEXT PRIMARY KEY,
    doc_id       TEXT NOT NULL REFERENCES documents(doc_id) ON DELETE CASCADE,
    chunk_index  INT  NOT NULL,
    heading      TEXT,                   -- the section the chunk came from
    content      TEXT NOT NULL,
    embedding    vector(1024) NOT NULL   -- Qwen3-Embedding-0.6B, same as the Vector DB guide
);

documents holds one row per source file with a checksum, so a re-run can skip files that have not changed. chunks holds the passages that are searched, each pointing back to its document.

Create two database roles, so the ingester can write while the query side only reads. Pick a password for each and substitute it for the placeholders, brackets included.

Still in psql:

CREATE ROLE rag_ingest LOGIN PASSWORD '<ingest-password>';
GRANT INSERT, UPDATE, DELETE, SELECT ON documents, chunks TO rag_ingest;

CREATE ROLE rag_query LOGIN PASSWORD '<query-password>';
GRANT SELECT ON documents, chunks TO rag_query;

Leave psql with \q.

Set the environment variables

The scripts read everything they need from the environment. Use the same two passwords as above; HF_TOKEN is the Hugging Face token; the two endpoints are the Endpoint Address shown on the embedder and rag-llm deployment pages, with or without the trailing /. If a password contains @, : or /, percent-encode it:

export VERDA_INFERENCE_KEY='<your-inference-api-key>'
export HF_TOKEN='<your-hugging-face-token>'
export EMBEDDER_ENDPOINT='<embedder-endpoint-address>'
export LLM_ENDPOINT='<rag-llm-endpoint-address>'
export VERDA_S3_ACCESS_KEY='<object-storage-access-key>'
export VERDA_S3_SECRET_KEY='<object-storage-secret-key>'
export RAG_BUCKET='rag-corpus-<name>'
export PG_INGEST_DSN='postgresql://rag_ingest:<ingest-password>@localhost/ragdb'
export PG_QUERY_DSN='postgresql://rag_query:<query-password>@localhost/ragdb'

These variables last only for the current shell. If you disconnect from the instance, run this block again after you reconnect.

Check the database connection before going further:

psql "$PG_INGEST_DSN" -c "select 1"

The expected output is a one-row table containing 1:

 ?column?
----------
        1
(1 row)

That means the rag_ingest role and password match. If you get password authentication failed, the DSN password differs from the one you set; reset it in psql as postgres on ragdb with ALTER ROLE rag_ingest PASSWORD '<ingest-password>'; and run the check again.

Step 4: Chunk and ingest

Install the Python packages in a virtual environment, in the same shell where you set the variables:

sudo apt-get install -y python3-venv
python3 -m venv ~/vecenv
source ~/vecenv/bin/activate
pip install requests "psycopg[binary]" numpy boto3 tokenizers

The ingester

The ingester splits each document into chunks, short passages that are embedded and searched individually.

Save the following script as ingest.py in your home directory:

ingest.py
# ingest.py
import hashlib, os, re, requests
import boto3, numpy as np, psycopg
from tokenizers import Tokenizer

BUCKET = os.environ["RAG_BUCKET"]
ENDPOINT = os.environ["EMBEDDER_ENDPOINT"].rstrip("/") + "/v1/embeddings"
KEY = "".join(os.environ["VERDA_INFERENCE_KEY"].split())   # strips stray whitespace
TARGET_TOKENS = 450        # chunk size; see "How chunking works"
BATCH = 64

tokenizer = Tokenizer.from_pretrained("Qwen/Qwen3-Embedding-0.6B")   # counts tokens like the model
FENCE = re.compile(r"^\s*(```|~~~)")
HEADING = re.compile(r"^#{1,6}\s+(.*)$")


def embed(texts: list[str]) -> np.ndarray:
    """The single source of truth for turning text into vectors. query.py imports it."""
    resp = requests.post(
        ENDPOINT,
        headers={"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"},
        json={"model": "embedder", "input": texts},
        timeout=300,               # generous: covers cold start
    )
    resp.raise_for_status()
    rows = [d["embedding"] for d in sorted(resp.json()["data"], key=lambda d: d["index"])]
    vecs = np.asarray(rows, dtype=np.float32)
    vecs /= np.linalg.norm(vecs, axis=1, keepdims=True)   # unit length, for cosine
    return vecs


def blocks(md: str):
    """Yield (heading, text) units: paragraphs, and fenced code blocks kept whole."""
    heading, buf, fence = "", [], None
    for line in md.split("\n") + [""]:
        if fence is None and FENCE.match(line):
            if buf:
                yield heading, "\n".join(buf); buf = []
            fence = FENCE.match(line).group(1); buf = [line]
        elif fence and line.strip().startswith(fence):
            buf.append(line); yield heading, "\n".join(buf); buf, fence = [], None
        elif fence:
            buf.append(line)
        elif HEADING.match(line):
            if buf:
                yield heading, "\n".join(buf); buf = []
            heading = HEADING.match(line).group(1).strip()
        elif line.strip() == "":
            if buf:
                yield heading, "\n".join(buf); buf = []
        else:
            buf.append(line)


def chunk(md: str):
    """Pack consecutive units under one heading into chunks of about TARGET_TOKENS."""
    chunks, cur, cur_heading, cur_tokens = [], [], None, 0
    for heading, text in blocks(md):
        n = len(tokenizer.encode(text).ids)
        if cur and (heading != cur_heading or cur_tokens + n > TARGET_TOKENS):
            chunks.append((cur_heading, "\n\n".join(cur))); cur, cur_tokens = [], 0
        cur.append(text); cur_heading = heading; cur_tokens += n
    if cur:
        chunks.append((cur_heading, "\n\n".join(cur)))
    return chunks


def ingest():
    s3 = boto3.client(
        "s3", endpoint_url="https://objects.fin-03.verda.storage", region_name="us-east-1",
        aws_access_key_id=os.environ["VERDA_S3_ACCESS_KEY"],
        aws_secret_access_key=os.environ["VERDA_S3_SECRET_KEY"],
    )
    conn = psycopg.connect(os.environ["PG_INGEST_DSN"])
    pages = s3.get_paginator("list_objects_v2").paginate(Bucket=BUCKET, Prefix="raw/")
    for obj in (o for p in pages for o in p.get("Contents", []) if o["Key"].endswith(".md")):
        key = obj["Key"]
        text = s3.get_object(Bucket=BUCKET, Key=key)["Body"].read().decode("utf-8")
        doc_id = key[len("raw/"):-len(".md")]
        checksum = hashlib.sha256(text.encode()).hexdigest()

        with conn.cursor() as cur:
            cur.execute("SELECT checksum FROM documents WHERE doc_id = %s", (doc_id,))
            if (row := cur.fetchone()) and row[0] == checksum:
                print(f"{doc_id}: unchanged"); continue

            title = next((m.group(1) for m in map(HEADING.match, text.split("\n")) if m), doc_id)
            cur.execute(
                """INSERT INTO documents (doc_id, title, checksum) VALUES (%s, %s, %s)
                   ON CONFLICT (doc_id) DO UPDATE
                   SET title = EXCLUDED.title, checksum = EXCLUDED.checksum, ingested_at = now()""",
                (doc_id, title, checksum),
            )
            cur.execute("DELETE FROM chunks WHERE doc_id = %s", (doc_id,))   # replace wholesale

            parts = chunk(text)
            for i in range(0, len(parts), BATCH):
                batch = parts[i:i + BATCH]
                vecs = embed([f"{h}\n\n{c}" if h else c for h, c in batch])
                cur.executemany(
                    """INSERT INTO chunks (chunk_id, doc_id, chunk_index, heading, content, embedding)
                       VALUES (%s, %s, %s, %s, %s, %s)""",
                    [(f"{doc_id}#{i + j}", doc_id, i + j, h, c, v.tolist())
                     for j, ((h, c), v) in enumerate(zip(batch, vecs))],
                )
            conn.commit()
            print(f"{doc_id}: {len(parts)} chunks")


if __name__ == "__main__":
    ingest()

Four functions do the work:

  • embed() turns text into vectors. It is the function from the Vector DB guide.
  • blocks() and chunk() split a Markdown file into chunks, one section at a time, keeping code blocks whole. How chunking works explains the rules and how to tune the size.
  • ingest() reads every .md file under raw/ in the bucket, skips files that have not changed since the last run, and replaces the chunks of any file that has.

Run the script:

python3 ingest.py

The output shows one line per file:

block-volumes: 4 chunks
pricing-and-billing: 3 chunks
serverless-containers: 3 chunks

The numbers match the sections in each file: block-volumes.md has an intro and three headings, so four chunks. The sections are short, so none needed splitting further.

Build the index

Now that the rows are loaded, we will create the index. With this index, Postgres finds the nearest vectors without comparing the question against every chunk. Building it after the load, rather than before, means the inserts did not have to update it as they went.

Run this as the postgres superuser:

sudo -u postgres psql -d ragdb
SET maintenance_work_mem = '2GB';

CREATE INDEX chunks_embedding_idx
    ON chunks USING hnsw (embedding vector_cosine_ops)
    WITH (m = 16, ef_construction = 64);

Leave psql with \q.

Step 5: Ask a question

Save this as query.py in the same directory as ingest.py, since it imports embed() from it. Keep the venv active and the environment variables from Step 3 set:

query.py
# query.py
import os, requests, psycopg
from ingest import embed

LLM = os.environ["LLM_ENDPOINT"].rstrip("/") + "/v1/chat/completions"
KEY = "".join(os.environ["VERDA_INFERENCE_KEY"].split())
conn = psycopg.connect(os.environ["PG_QUERY_DSN"])

SQL = """
SELECT c.heading, c.content, d.title, d.doc_id,
       1 - (c.embedding <=> %(q)s::vector) AS similarity
FROM chunks c JOIN documents d ON d.doc_id = c.doc_id
ORDER BY c.embedding <=> %(q)s::vector
LIMIT %(k)s
"""

SYSTEM = """Answer the question using only the numbered sources below.
Put the source number in square brackets after each claim, like [2].
If the sources do not contain the answer, say so. Do not invent source numbers."""


def retrieve(question: str, k: int = 6):
    qvec = embed([question])[0].tolist()
    with conn.cursor() as cur:
        cur.execute("SET LOCAL hnsw.ef_search = 100")
        cur.execute(SQL, {"q": qvec, "k": k})
        cols = [d.name for d in cur.description]
        return [dict(zip(cols, r)) for r in cur.fetchall()]


def answer(question: str):
    hits = retrieve(question)
    sources = "\n\n".join(f"[{i}] {h['title']} — {h['heading']}\n{h['content']}"
                          for i, h in enumerate(hits, 1))
    resp = requests.post(
        LLM,
        headers={"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"},
        json={"model": "rag-llm", "temperature": 0.2,
              "chat_template_kwargs": {"enable_thinking": False},   # no <think> preamble
              "messages": [{"role": "system", "content": SYSTEM},
                           {"role": "user", "content": f"Sources:\n\n{sources}\n\nQuestion: {question}"}]},
        timeout=300,
    )
    resp.raise_for_status()
    return resp.json()["choices"][0]["message"]["content"], hits


if __name__ == "__main__":
    import sys
    text, hits = answer(" ".join(sys.argv[1:]) or "Am I charged for a volume when the instance is shut down?")
    print(text, "\nSources")
    for i, h in enumerate(hits, 1):
        print(f"  [{i}] {h['title']} — {h['heading']}  ({h['doc_id']}.md, {h['similarity']:.2f})")

retrieve() is the search from the Vector DB guide, joined to documents for the title. answer() numbers the retrieved chunks, tells the model to cite by number, and prints the numbered list itself, so a citation always resolves to a real chunk.

Run the script with a question:

python3 query.py "Am I charged for a volume when the instance is shut down?"

On the sample corpus, the output is:

Yes, you are charged for a volume when the instance is shut down. A Block Volume is charged
whether or not it is attached to an instance, and whether or not that instance is running [2].
Sources
  [1] Pricing and billing — Shutting down and deleting instances  (pricing-and-billing.md, 0.75)
  [2] Pricing and billing — Storage charges  (pricing-and-billing.md, 0.72)
  [3] Pricing and billing — Pricing and billing  (pricing-and-billing.md, 0.61)
  [4] Block Volumes — Attaching and resizing  (block-volumes.md, 0.61)
  [5] Block Volumes — Deleting a volume  (block-volumes.md, 0.59)
  [6] Block Volumes — Cloning a volume  (block-volumes.md, 0.56)

The claim points at a section of one of the three files you created. All six retrieved chunks are listed with their similarity to the question; the model used only the one it needed.

Now ask something the documents do not cover:

python3 query.py "Am I a cat?"
None of the provided sources address the question "Am I a cat?" Therefore, the answer cannot
be determined from the given sources.
Sources
  [1] Serverless Containers — Serverless Containers  (serverless-containers.md, 0.24)
  ...

The model says so instead of guessing, because the system prompt tells it to answer from the sources only. The similarity scores, around 0.2 instead of 0.7, show that nothing relevant was retrieved.

The score after each source runs from 0 to 1:

  • 1 means the chunk has the same meaning as the question.
  • 0 means the chunk is unrelated to it.

Congratulations! You now have a working RAG service.


After it works

When an answer is wrong

A citation tells you which source to check; it does not guarantee the model read that source correctly. Look at the Sources list first. If the passage that answers the question is not among the six, the model never saw it, and the fix is in chunking or retrieval rather than in the prompt.

How chunking works

The embedding model turns a piece of text into one vector that stands for its meaning as a whole, and search returns the vectors closest to the question's. A vector for a whole page captures what the page is about, not the details inside it; a vector for one section does capture that section's details, so it matches a question about them. Documents are therefore split into chunks of a few hundred tokens, each embedded on its own, and it is chunks that are searched and handed to the LLM.

Two rules matter when your documents are Markdown files with headings and text blocks, as most product docs are, and ingest.py applies both:

  • Split on headings, so every chunk carries its section heading. The heading helps retrieval and is what the citation shows.
  • Never cut a command or code block in two. It is added to a chunk whole or not at all, because half a command looks like a complete answer but fails when run.

The chunk size is TARGET_TOKENS at the top of ingest.py. 450 suits documentation pages: smaller chunks are more precise but lose surrounding context, larger ones keep context but match specific questions less well. Go lower for FAQ-style content, higher for long narrative text, and re-run ingest.py on the whole corpus after any change.

Updating the corpus

Edit or add Markdown files in corpus/, then run verda oss sync and python3 ingest.py again. Files that have not changed (their checksum matches the one stored in documents) are skipped, so only what changed is embedded.


Conclusion

You now have documents in Object Storage, their passages indexed in pgvector, an embedding model and an LLM on Serverless Containers, and a script that turns a question into a cited answer.

To clean up. The instance, the volume and the bucket keep billing while they exist. When you are finished, discontinue the instance, permanently delete the volume, delete both container deployments, and remove the bucket together with its files:

verda oss rb s3://rag-corpus-<name> --force

--force deletes everything in the bucket before removing it, and asks for confirmation first.