Tomefile · local RAG · format 1.3 · forever free
Fine-tuning will not teach it your documents.
A local model can write. It cannot be trusted to remember a handbook, a policy, or a door code. Those are facts you look up. Fine-tune it for tone if you want. Put the facts and the embedder into one file, a .tome, and search them on 127.0.0.1.
Local RAGOne fileFree forever
- id
- com.example.offices
- revision
- 1
- bytes
- 43150328
- signature
- unsigned
- embedding
- BAAI/bge-small-en-v1.5
- dimensions
- 384
- metric
- cosine
- embed
- models/embed.gguf
- chunks
- 3072
- sources
- 1024
- characters
- 1154762
- top_k
- 3
- knowledge
- internal attribution=true
- builder
- tomefile 0.1.0
Two good tools. Two different jobs.
People take a small instruct model, fine-tune it on the company wiki, and expect “how many vacation days?” to come back exact. The demo sounds good. The fact is not in the model. Retrieval is the right idea. The RAG most people build still needs the internet.
Fine-tune the voice
Tone, format, company style, how it should act. That is what fine-tuning is for. Keep it. Use a .tome for the facts the model should not have to remember.
RAG, usually not local
An embeddings API, a hosted vector store, tool calls, an account. Retrieval works. The question and the index still leave your machine.
RAG as a local file
Chunks, vectors, and the embedder packed into one file. The chat model you already run. The port is 127.0.0.1. No extra API calls. No remote index.
Retrieval beat unsupervised fine-tuning for adding knowledge (Ovadia et al., 2024). Fine-tuning on new facts makes models learn them slowly and invent more (Gekhman et al., 2024). That is about facts, not about tone. More, with numbers →
Your model. Your documents. On your machine.
Tomefile does not replace Ollama, llama.cpp, or Open WebUI. It is the retrieval index those tools query locally. Typical RAG makes many API calls to a store you do not run. This one is a file.
Build
Chunk and embed a folder of .txt and .md. The file carries the embedder, or pins it by hash.
$ tomefile build docs/ \
--id com.acme.handbook \
--revision 1 \
--model bge-small-en-v1.5-q8_0.gguf \
--embed-id BAAI/bge-small-en-v1.5 \
--embed-license mit \
--output handbook.tome
Open locally
Use the chat model you already run. The embedder comes from the file. The port is 127.0.0.1.
$ tomefile serve handbook.tome \
--engine ollama \
--ollama-model llama3.1
http://127.0.0.1:8080/v1
Point any client
Open WebUI, Continue, a script, or curl. Each answer names the source file and character offsets.
$ curl http://127.0.0.1:8080/v1/chat/completions \
-H 'content-type: application/json' \
-d '{"messages":[{"role":"user",
"content":"How many vacation days?"}]}'
The answer cites the passage. Not the weights.
The server finds the best chunks, puts them in the prompt, and returns what the local model said, plus tomefile.sources.
{
"choices": [{ "message": {
"role": "assistant",
"content": "You get 18 vacation days per calendar year."
}}],
"tomefile": {
"id": "com.example.handbook",
"revision": 1,
"sources": [
{ "source": "leave.md", "start": 0, "end": 243 },
{ "source": "onboarding.md", "start": 0, "end": 259 },
{ "source": "kitchen.txt", "start": 0, "end": 286 }
],
"license": "internal"
}
}
without the file. Qwen2.5-1.5B, temperature 0, a private handbook it had never seen. Sounds sure. Wrong.
through the file, same model. Same as giving it the correct paragraph. Retrieval did the work.
Made-up handbook, Apple M1. The corpus is easy to search. The test is whether the facts reach the model, not a retrieval contest. Paper →
Knowledge and models, into one file.
A .tome packs the documents, the vectors, and the embedder. A .tbox can add the chat model too. Checksums cover every member. Sign it with Ed25519 if you publish. Copy that one file to another machine and open it. You can check it was not changed.
sha256sum format.sqlite-vec vectors.Opening a 3,072-chunk file and retrieving the first context took 0.58 s on a clean machine; rebuilding that index took 63.7 s. How to pack a .tbox is in Get started. Read the format →
Local. Free forever.
No account. No API key. The server binds 127.0.0.1. The format is an open specification; the files you build open without us.
- Runs next to Ollama or llama.cpp
- No calls to the internet
- Readable with unzip and sqlite3
- Free to implement, with no royalties