Kremis

A knowledge graph that records, associates, and retrieves — but never invents. Same input, same output. Honest about what is fact, what is absent, and nothing in between.

The problem

An LLM is fluent and confident even when it is wrong. Ask it for a fact it never saw and it will often produce a plausible one anyway — no trace, no provenance, no way to tell recall from invention.

The usual framing says the hard problem is generating knowledge. Kremis takes the opposite posture: the hard problem is being able to verify it. A model is a good narrator and a poor ledger. So give it a ledger it cannot lie to.

The approach

Kremis is the substrate, not the model. It holds what is actually known and answers only from that — a missing entity returns an explicit not found, never a guess.

The proof

A reproducible fabrication benchmark ships with the repo. It builds a closed registry of 9 fictional services joined by 5 one-way dependencies, then asks 24 questions of the form "does A depend on B, directly or transitively?" — 8 have an answer, 16 do not, and no answer exists for them anywhere. Nothing in the prompt invites invention: the facts are supplied and UNKNOWN is offered.

Against qwen3.5:4b at temperature 0, over 5 runs:

System False assertions Answer accuracy
Kremis (/query + /certify) 0.00 % 100 %
LLM holding the whole registry 0.00 % 100 %
LLM + naive retrieval 0.00 % 75 %
LLM, no context 0.00 % 0 %

I won't oversell that zero. On a world this small a capable model doesn't fabricate either — given every fact it needs, qwen3.5:4b matches the substrate. But capability isn't guaranteed: phi4-mini, another current local 4B, holds the identical registry and still asserts a dependency in reverse on every run. Which model you happen to run decides it.

Kremis's zero is structural, not measured. Dependencies are one-way edges, so a reversed path isn't there to be found: the query returns grounding: "unknown" and /certify issues a certificate carrying no evidence, bound to a BLAKE3 hash of the graph state. Nor is this a like-for-like race — the model reads English and has to resolve the services itself, while Kremis is handed the ids. A graph that cannot represent an edge cannot invent one, and saying so proves little. What isn't free is the certificate: an absence bound to a hash, checkable by someone who doesn't trust the system that issued it.

# ingest signals, then serve
cargo run -p kremis -- init
cargo run -p kremis -- ingest -f examples/sample_signals.json -t json
cargo run -p kremis -- server

# the benchmark — with a local model, or Kremis alone
python benchmark/run.py --model qwen3.5:4b --runs 5
python benchmark/run.py --skip-llm

Status

Alpha at v0.21.4 — functional and tested; breaking changes may still land before v1.0, so pin your version. The core is published on crates.io as kremis-core, there's a Docker image, and the full reference lives at kremis.mintlify.app. Rust throughout, Apache 2.0, built with AI as an ally and a deterministic spine.

The project's spine

Keep it minimal. Keep it deterministic. Keep it grounded. Keep it honest.