tomefile.
Why · local RAG

Fine-tune the voice. Retrieve the facts.

Fine-tuning is a good tool. RAG is a good tool. Most setups put the handbook in the wrong place, then call an embeddings API and a hosted index. The local version is a file next to the model you already run.

01 · Weights

Fine-tuning is for how it talks. Keep it.

A LoRA is a good way to give a local model a company voice: how it greets, how it formats, what it refuses, which words it prefers. Do that. Then let it read. The mistake is putting the wiki into the same training and expecting “how many vacation days?” to come back as the exact number. The model has learned a tone. It has not stored the number.

In tests, retrieval beat unsupervised fine-tuning for adding knowledge (Ovadia et al., 2024). Fine-tuning on new facts makes models learn them slowly and invent more (Gekhman et al., 2024). That is about facts. Tone is a different job. The usual mix is a small fine-tune for how it talks, and a .tome for the documents. You can cite a passage, fix one paragraph, and keep the LoRA.

In our test, Qwen2.5-1.5B answered 0 of 256 questions about a private handbook on its own, and 256 of 256 through the file. Same model. The index did the remembering. A style fine-tune of that model would still need the file for the facts.

02 · Local

RAG is right. Most RAG is not local.

Retrieval is how you give a model documents it should not have to remember. What people actually build is rarely on the machine: an embeddings HTTP call, a hosted vector database, an agent that “uses tools,” an API key. The answer can still use the documents. The question left the laptop. The index is someone else’s service.

Ollama, LM Studio, and llama.cpp already run the chat model locally. Open WebUI, Continue, and GPT4All already speak the OpenAI API. Tomefile is the retrieval index those tools query on 127.0.0.1. The embedder is inside the file, so the vectors match. No call to a remote store. The question stays on your machine.

Sending the same file to a colleague is a bonus of putting the index in a file. It is not the main reason to build it.

03 · Servers

When to use a hosted vector database instead.

pgvector, Qdrant, and similar tools run a live shared index: many people write, pages change all the time, one server everyone uses. If that is your problem, use them. A .tome is a finished version you query on your machine.

Hosted vector storeTomefile
JobA live index many people writeA local RAG index you open
You operateA server: uptime, backups, accessNothing at query time. A file.
The question goesTo that serverNowhere. 127.0.0.1
The embedderYour app must remember itPinned by SHA-256 in the file
UpdatesLive writesA new revision under the same id
Best forData that changes constantlyA corpus you can point a local model at
04 · Apps

Not a new chat app.

AnythingLLM, Open WebUI, and GPT4All already let you upload documents and chat. Their index stays in app data. Tomefile does not replace them. Point one at http://127.0.0.1:8080/v1 and they query a .tome they did not have to load again.

05 · Fit

Where it fits, honestly.

A local model and a handbook

You already run Ollama. You have policies in Markdown. You want cited answers. A LoRA is optional, for voice.

Docs next to a release

Ship a signed .tome or .tbox with the version. One file: knowledge, embedder, and if you pack it, the chat model. Users check it and ask questions offline.

Air-gapped machines

Copy one verifiable file in. Bring llama-server once. Each corpus after that is another file, not another pipeline.

A colleague without the pipeline

They skip chunking and embedding. Useful. Not the main reason you built the file.

A LoRA plus a .tome

Fine-tune for style, format, how it should act. Retrieve the handbook. Use both. It is not one or the other.

Text that must stay secret

Chunk text is readable SQLite. Anyone holding the file can read the corpus.

Data that changes every minute

Use a live vector store. A file is a revision, not a stream.

06 · Paper

The long version.

Tomefile: A Portable File Format for Prebuilt Local Retrieval Indexes. Simeon Emanuilov, Faculty of Mathematics and Informatics, Sofia University. Why facts do not belong in the model weights, the format, and answers with and without the file.