tomefile.
Guide · reference CLI 0.1.0

From a folder of docs to local RAG.

Install, build, serve, point the client you already use at 127.0.0.1. Signing, revisions, and packing a chat model come after, when you need them.

00 · Needs

What you need.

  • Python 3.10 or newer.
  • llama-server from llama.cpp on your PATH. On macOS: brew install llama.cpp.
  • To build: an embedding model in GGUF, for example bge-small-en-v1.5-q8_0.gguf (37 MB). The receiver does not need it; it travels in the file.
  • To chat: a running Ollama, or any chat model in GGUF.
01 · Install

Install the CLI.

$ git clone https://github.com/s-emanuilov/tomefile.git
$ cd tomefile
$ python -m venv venv && source venv/bin/activate
$ pip install -e .

This adds one command, tomefile, with five subcommands: build, pack, graph, inspect, and serve.

02 · Build

Build a .tome from a folder.

$ tomefile build docs/ \
  --id com.acme.handbook \
  --revision 1 \
  --model bge-small-en-v1.5-q8_0.gguf \
  --embed-id BAAI/bge-small-en-v1.5 \
  --embed-license mit \
  --output handbook.tome
  • docs/ is read recursively. Version 1.3 takes .txt and .md in UTF-8 and skips other files.
  • --id is the package name that stays the same across rebuilds. Use a domain you control, reversed.
  • --embed-id names the model that produced the vectors. It must match the GGUF.
  • The knowledge license defaults to internal, not commercial, not redistributable. Change it with --knowledge-license, --knowledge-commercial, and --knowledge-redistribute.

Text is cut into 500-character chunks at whitespace, without overlap, and the receiver gets the top 3 per question.

03 · Serve

Open it on the receiving machine.

With Ollama already running:

$ tomefile serve handbook.tome --engine ollama --ollama-model llama3.1

Or with a chat GGUF:

$ tomefile serve handbook.tome --llm qwen2.5-1.5b-instruct-q4_k_m.gguf

Before it opens a port, serve verifies every checksum and any signature, checks the embedder hash, and starts the embedder from the file. It then prints the base URL, http://127.0.0.1:8080/v1 by default. It listens only on 127.0.0.1; set another port with --port.

04 · Connect

Use the app you already have.

Any OpenAI-compatible client works. Set the base URL to http://127.0.0.1:8080/v1. The gateway does not check API keys, so any value works where a client demands one. It serves the loaded file whatever model you send; GET /v1/models returns the package name for clients that ask first.

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="local")
reply = client.chat.completions.create(
    model="handbook",
    messages=[{"role": "user", "content": "How many vacation days do I get?"}],
)
print(reply.choices[0].message.content)

Every response, including each streamed chunk, carries a tomefile object with the package id, revision, and the source, start, and end of each chunk used. Streaming and temperature pass through.

05 · Inspect

See what is inside.

$ tomefile inspect handbook.tome
$ tomefile inspect handbook.tome --chunks

inspect verifies the checksums and prints the package id, revision, signature, embedder, chunk and source counts, member sizes, and one line per source file. It starts no model. --chunks also prints every chunk with its source and character offsets.

06 · Sign

Sign what you publish.

A signature lets a receiver check that the file came from you and was not changed. Create a key once and keep it private:

$ openssl genpkey -algorithm ed25519 -out publisher.pem
$ tomefile build docs/ ... --sign publisher.pem --signer com.acme

This needs OpenSSL 1.1.1 or newer. The LibreSSL that ships with macOS has no Ed25519; brew install openssl provides one that does.

Publish your public key where receivers already trust you, such as your website. This prints it in the form the CLI expects:

$ openssl pkey -in publisher.pem -pubout -outform DER | tail -c 32 | base64

The receiver saves that line as ~/.config/tomefile/trusted/com.acme and serves with --require-signature. Unsigned files, and files signed by any other key, are refused.

pack and graph rewrite the checksums, so they drop an input signature. Pass --sign again.
07 · Update

Ship revision 2.

Rebuild with the same --id and a higher --revision. The receiver replaces one file.

$ tomefile build docs/ --id com.acme.handbook --revision 2 ... --output handbook.tome

Your first build put the embedder in your model cache, ~/.cache/tomefile/models/, so later builds reference it by hash instead of copying it. Receivers who opened revision 1 already have it in theirs. For someone who never did, add --embed-model and the file carries the weights again.

08 · Pack

Optional: pack the chat model into the same file.

A .tbox is knowledge, embedder, and chat model in one file you can copy and check. By default it stores a SHA-256 reference to the chat weights, not a copy. Add --sign so the receiver can verify the publisher key they already trust.

$ tomefile pack handbook.tome \
  --llm qwen2.5-1.5b-instruct-q4_k_m.gguf \
  --chat-license apache-2.0 \
  --output handbook.tbox

Add --embed --redistributable only when the model license allows shipping the weights inside the file.