auto provider
works without any setup; this page is mainly for tuning, switching to local,
or controlling cost.
Embeddings are what power the “meaning matching” side of memory search. They
convert text into a numerical representation that lets Comis find memories based
on meaning, not just matching words. You do not need to understand the math —
just choose a provider and Comis handles the rest.
What is wired today
Two embedding implementations ship inpackages/memory:
- OpenAI embeddings (
embedding-provider-openai.ts) — calls thetext-embedding-3-smallmodel by default (1536 dimensions). - Local embeddings (
embedding-provider-local.ts) — runs a small transformer model entirely on-device using a downloaded GGUF file.
packages/memory/src/embedding-provider-factory.ts chooses
between them based on the embedding.provider setting. The default
"auto" mode tries local first and falls back to OpenAI if loading fails.
Why embeddings matter
Without embeddings, memory search only finds exact word matches. With embeddings, your agent can find related memories even when different words are used. For example:- A memory about “cooking pasta” would match a search for “making spaghetti”
- A memory about “budget planning” would match a search for “financial forecasting”
- A memory about “fixing the deployment” would match a search for “resolving the release issue”
Embedder vs reranker
Comis runs two different local GGUF models (both vianode-llama-cpp) at two
different stages of recall — it is easy to confuse them, so the distinction
matters:
The embedder is a bi-encoder: it encodes the query and each memory
separately into vectors, and similarity is the distance between them — fast,
and what powers the meaning-matching search lane. The
reranker is a cross-encoder: it processes the query and a candidate
jointly for a sharper relevance judgement, but is more expensive, so it only
re-scores the handful of candidates that already passed fusion, and it is
opt-in.
Choosing a provider
Comis supports three embedding provider options: auto (default) — Tries local first, then falls back to remote. This is the best option for most users because it works without any setup. Comis automatically downloads a local model and uses it. If the local model fails to load, it falls back to the OpenAI API. local — Runs entirely on your machine using a downloaded model. This option is free, private (no data leaves your machine), and works offline. It requires approximately 500MB of disk space for the model and uses some RAM while running. GPU acceleration is used automatically when available. openai — Uses OpenAI’s embedding API. This option is fast and accurate but costs a small amount per search (fractions of a cent) and requires an OpenAI API key. Choose this if you want the highest quality embeddings or if your machine does not have enough resources for the local model.Multilingual deployments
The default on-device model (nomic-embed-text-v1.5) is English-centric. On a
deployment that mostly handles non-Latin text (Hebrew, Arabic, CJK, Cyrillic), its
semantic recall degrades: lexical full-text search and the multilingual reranker
(bge-reranker-v2-m3) still carry recall, but meaning-based matches across
paraphrase and synonyms weaken. For those deployments, switch to a multilingual
embedder:
- On-device
bge-m3— private, $0, ~635MB download, 1024-d. Setembedding.local.modelUrito a bge-m3 GGUF (e.g.hf:gpustack/bge-m3-GGUF:bge-m3-Q8_0.gguf) andembedding.multilingual: true. - OpenAI
text-embedding-3-small— hosted, multilingual, 1536-d, but sends memory text to OpenAI and costs per embed. Setembedding.provider: openaiandembedding.multilingual: true.
comis init wizard asks about this at install (“mostly non-English?”) and
writes the right embedding.* block for you. Non-interactive installs use
--embedding-multilingual with --embedding-provider local|openai (and
--embedding-api-key for OpenAI when the main provider is not OpenAI). Changing
the embedder later re-embeds existing memories automatically (embedding.autoReindex).
embedding.multilingual is an advisory flag — it declares the embedder
multilingual for the comis fleet model-health line and gates no behavior. The
embedder itself is selected by embedding.provider / embedding.local.modelUri.Configuration
Local provider example
~/.comis/config.yaml
OpenAI provider example
~/.comis/config.yaml
Persistent caching
Comis uses a two-tier embedding cache to avoid redundant API calls:- L1 (in-memory LRU) — Fast access for recent embeddings. Controlled by
embedding.cache.maxEntries. Lost on daemon restart. - L2 (SQLite-backed persistent) — Durable storage that survives daemon restarts. Embeddings computed once are reused across sessions indefinitely.
~/.comis/config.yaml
If you change your embedding provider (for example, switching from local to
OpenAI), Comis automatically re-indexes all existing memories to use the new
model. This happens in the background and does not interrupt normal operation.
You can disable this with
autoReindex: false if you prefer to manage
re-indexing manually.Re-embedding strategy
Switching providers (or models) means existing embeddings no longer match new ones — they live in a different vector space. Comis tracks the embedding fingerprint (provider + model + dimensions) and detects mismatches at startup. When the new model’s dimensions differ from the old one’s (for example, switching from a local 768-dimension model to OpenAI’s 1536-dimensiontext-embedding-3-small), the vector tables themselves are rebuilt at startup —
a vector table’s dimension is fixed when the table is created, so Comis drops
and recreates it at the new dimension and queues every memory for
re-embedding. The rebuild is logged at INFO (Vector table rebuilt for new embedding dimension) and recorded in the boot model_health diagnostic, so
comis fleet drill-downs can confirm it ran.
If the vector lane ever fails at query time (for example, a vector-table /
embedder mismatch while the daemon is running), recall degrades to
FTS-only for that query instead of failing outright — the degradation is
recorded in the session trajectory (memory.recall_degraded), surfaced in
comis explain’s recall section, and counted as a
health_signal:recall_degraded finding in comis fleet.
When embedding.autoReindex: true (the default), a background batch indexer
re-embeds every existing memory in batches of embedding.batch.batchSize
(default 100) until the database is consistent again. The agent stays
responsive throughout — re-indexing runs out of band, and search keeps using
whatever embeddings are valid at any given moment.
The fingerprint manager
(packages/memory/src/embedding-fingerprint.ts) and batch indexer
(packages/memory/src/embedding-batch-indexer.ts) handle this automatically.
Cost note
OpenAI’stext-embedding-3-small is cheap: roughly $0.02 per 1M tokens
of input. A typical memory entry is 50–200 tokens, so embedding 10,000
memories costs around $0.02–$0.04 total — and once cached (see
“Persistent caching” above), re-indexing is free.
The local provider has zero per-call cost but uses ~500 MB of disk and some
RAM while running.
Disabling embeddings
If you do not want to use meaning-based search at all, setembedding.enabled
to false. Memory search will still work using text matching only. This is
useful if you want to minimize resource usage or if your agent’s memory is small
enough that text matching is sufficient.
Search
How text matching and meaning matching work together.
Memory
The different types of memories your agent stores.
