From a folder of docs to local RAG.
Install, build, serve, point the client you already use at 127.0.0.1. Signing, revisions, and packing a chat model come after, when you need them.
What you need.
- Python 3.10 or newer.
llama-serverfrom llama.cpp on yourPATH. On macOS:brew install llama.cpp.- To build: an embedding model in GGUF, for example bge-small-en-v1.5-q8_0.gguf (37 MB). The receiver does not need it; it travels in the file.
- To chat: a running Ollama, or any chat model in GGUF.
Install the CLI.
$ git clone https://github.com/s-emanuilov/tomefile.git
$ cd tomefile
$ python -m venv venv && source venv/bin/activate
$ pip install -e .
This adds one command, tomefile, with five subcommands: build, pack, graph, inspect, and serve.
Build a .tome from a folder.
$ tomefile build docs/ \
--id com.acme.handbook \
--revision 1 \
--model bge-small-en-v1.5-q8_0.gguf \
--embed-id BAAI/bge-small-en-v1.5 \
--embed-license mit \
--output handbook.tome
docs/is read recursively. Version 1.3 takes.txtand.mdin UTF-8 and skips other files.--idis the package name that stays the same across rebuilds. Use a domain you control, reversed.--embed-idnames the model that produced the vectors. It must match the GGUF.- The knowledge license defaults to
internal, not commercial, not redistributable. Change it with--knowledge-license,--knowledge-commercial, and--knowledge-redistribute.
Text is cut into 500-character chunks at whitespace, without overlap, and the receiver gets the top 3 per question.
Open it on the receiving machine.
With Ollama already running:
$ tomefile serve handbook.tome --engine ollama --ollama-model llama3.1
Or with a chat GGUF:
$ tomefile serve handbook.tome --llm qwen2.5-1.5b-instruct-q4_k_m.gguf
Before it opens a port, serve verifies every checksum and any signature, checks the embedder hash, and starts the embedder from the file. It then prints the base URL, http://127.0.0.1:8080/v1 by default. It listens only on 127.0.0.1; set another port with --port.
Use the app you already have.
Any OpenAI-compatible client works. Set the base URL to http://127.0.0.1:8080/v1. The gateway does not check API keys, so any value works where a client demands one. It serves the loaded file whatever model you send; GET /v1/models returns the package name for clients that ask first.
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="local")
reply = client.chat.completions.create(
model="handbook",
messages=[{"role": "user", "content": "How many vacation days do I get?"}],
)
print(reply.choices[0].message.content)
Every response, including each streamed chunk, carries a tomefile object with the package id, revision, and the source, start, and end of each chunk used. Streaming and temperature pass through.
See what is inside.
$ tomefile inspect handbook.tome
$ tomefile inspect handbook.tome --chunks
inspect verifies the checksums and prints the package id, revision, signature, embedder, chunk and source counts, member sizes, and one line per source file. It starts no model. --chunks also prints every chunk with its source and character offsets.
Sign what you publish.
A signature lets a receiver check that the file came from you and was not changed. Create a key once and keep it private:
$ openssl genpkey -algorithm ed25519 -out publisher.pem
$ tomefile build docs/ ... --sign publisher.pem --signer com.acme
This needs OpenSSL 1.1.1 or newer. The LibreSSL that ships with macOS has no Ed25519; brew install openssl provides one that does.
Publish your public key where receivers already trust you, such as your website. This prints it in the form the CLI expects:
$ openssl pkey -in publisher.pem -pubout -outform DER | tail -c 32 | base64
The receiver saves that line as ~/.config/tomefile/trusted/com.acme and serves with --require-signature. Unsigned files, and files signed by any other key, are refused.
pack and graph rewrite the checksums, so they drop an input signature. Pass --sign again.Ship revision 2.
Rebuild with the same --id and a higher --revision. The receiver replaces one file.
$ tomefile build docs/ --id com.acme.handbook --revision 2 ... --output handbook.tome
Your first build put the embedder in your model cache, ~/.cache/tomefile/models/, so later builds reference it by hash instead of copying it. Receivers who opened revision 1 already have it in theirs. For someone who never did, add --embed-model and the file carries the weights again.
Optional: pack the chat model into the same file.
A .tbox is knowledge, embedder, and chat model in one file you can copy and check. By default it stores a SHA-256 reference to the chat weights, not a copy. Add --sign so the receiver can verify the publisher key they already trust.
$ tomefile pack handbook.tome \
--llm qwen2.5-1.5b-instruct-q4_k_m.gguf \
--chat-license apache-2.0 \
--output handbook.tbox
Add --embed --redistributable only when the model license allows shipping the weights inside the file.