The complete tour

How Centauri works

Everything under the hood — the storage engine that never erases, the bi-temporal data model, the CeQL query language, how Centauri talks to local and cloud AI, and honest comparisons to the databases you already know. If the showcase is the what, this is the how.

1 · The big picture

Centauri is one self-contained binary that is two things at once: an immutable system of record and a private AI engine that reasons over it — with zero third-party dependencies.

Ordinary databases store the latest value and overwrite the rest. Centauri stores every fact, forever, on an append-only log. An "update" appends a superseding fact; a "delete" appends a RETIRE marker. Because the past is never destroyed, Centauri can answer questions no ordinary store can — what was true at any moment, what you believed then, why it changed, and whether anyone tampered with it — and it keeps the same discipline for AI: every model answer and every human correction is itself a fact.

Clients CeQL · SQL · REST psql / JDBC / BI AI agents (MCP) Python SDK Centauri — one binary Write path (write-then-apply) append hash-chain fsync apply→RAM Storage tiers hot logJSONL, appended sealed segmentscompressed + zone maps cold tierlocal / S3 In-RAM indexes current fact per subject · vectors secondary & zone-map index Query engine CeQL · time-travel · causal BM25 + vector search RAG / ENRICH / ASK Replay rebuilds every in-RAM index from the log — deterministically. Local LLM Ollama / LocalAI OpenAI-compatible nothing leaves the box AI (optional)

2 · The atom: a fact

The unit of storage is not a row or a document — it's an immutable fact (event). Each fact carries who/what it's about, the value, two clocks, and trust metadata:

  • Subject — what the fact is about, e.g. item:123456/store:4412.
  • Facet — which "reality" the fact belongs to (source, register, shelf, …), so the same subject can hold several parallel truths.
  • Value — the payload (a flexible field map), optionally schema-checked.
  • Two clocks — valid time (when it was true in the world) and transaction time (when Centauri learned it). This is what makes it bi-temporal.
  • Provenance & confidence — where the fact came from and how much to trust it.
  • Causal links — directed edges (TRIGGERED, SUPERSEDES, CORRECTS, ENRICHED_FROM, …) so lineage is first-class data, not a join you invent later.
Updates and deletes are also facts. "Change the price" appends a new fact that supersedes the old one; "remove this" appends a RETIRE marker. The original bytes are never modified — that's the whole point.

3 · The storage engine

A hot append-only log, sealed into compressed segments, tiered to cheap cold storage — with the working set indexed in RAM.

Write-then-apply

A write is appended to the log, linked into the hash chain, and made durable (fsync) before it changes any in-memory state. If the process crashes mid-write, recovery replays the log and lands in a consistent state — memory never gets ahead of disk. An optional group commit coalesces concurrent writes into a single fsync for throughput.

Hot log → sealed segments → cold tier

Recent facts live in a hot JSONL log. Once it grows past a threshold (or on demand) the tail is sealed into a segment that is flate-compressed (typically 5–10× on cold data) and carries uncompressed zone-map statistics + a Merkle root in its manifest — so the engine can prune and verify a segment without decompressing it. Sealed segments can be pushed to an S3-compatible cold tier (AWS/MinIO/R2/B2) over stdlib request signing; segments fetched from untrusted storage are Merkle-verified before use and LRU-cached.

Replay & determinism

Every in-RAM index (current-fact pointers, vectors, secondary indexes) is rebuilt by replaying the log through one apply step, so a restart reconstructs exactly the same state. A periodic checkpoint means open/recovery replays only the tail since the last checkpoint, not the whole history — recovery time stays bounded no matter how much data accumulates.

Crypto-erasure

For "right to be forgotten" without breaking the audit trail: a segment's payloads are encrypted with a per-segment AES-256-GCM key; destroying that key makes the payloads permanently unreadable while the hash chain — and therefore the proof of history — stays intact.

4 · Tamper-evidence (without a blockchain)

Every committed line is linked into a SHA-256 hash chain: each entry's hash folds in the previous entry's hash, so changing any past byte breaks every hash after it. Each sealed segment also publishes a Merkle root. The same chain is computed identically on the write path, on replay, and on raw ingestion (replication), and verify (or the dashboard's Verify button / /v1/verify) recomputes both to prove nothing was altered — with no downtime and no third party.

This gives ledger-grade integrity for audit (prove what was true, what you knew, and that nothing changed) — the kind of evidence SOX/HIPAA-style reviews ask for — without the cost and complexity of an actual blockchain.

5 · Bi-temporal & causal queries

Two clocks unlock the questions ordinary databases can't answer:

QuestionHow
What was true at a moment?AS OF '2026-03-15' — valid-time travel
What did we believe at a moment?AS KNOWN AT '2026-03-01' — transaction-time travel (the audit superpower)
Why did it change?WHY <event> / EFFECTS <event> — walk the causal graph
Did a change actually land?PENDING — find facts distributed but never activated
Where do sources disagree?DISAGREE ON <field> — cross-facet conflict detection

Causal links make lineage queryable directly: MATCH item:* CAUSES order:* finds chains across subjects without hand-written joins.

6 · The CeQL query language

CeQL is the native language where time, cause, and trust are syntax — not bolted-on functions. It comes in three interchangeable forms (text, JSON-AST, and REST) and has a plain-English translator for the rest.

Reading & time travel

FACTS OF item:* WHERE price_cents > 700
HISTORY OF item:1                       -- the whole story of one subject
FACTS OF item:1 AS OF YESTERDAY          -- what was true then
FACTS OF item:1 AS KNOWN AT '2026-03-01'   -- what we believed then

Causality

WHY ev_01H...           -- what caused this fact
EFFECTS ev_01H...       -- what this fact caused
MATCH item:* CAUSES order:*   -- causal chains across subjects

Search — keyword, vector, hybrid

SEARCH 'late markdown' OF item:*            -- ranked BM25, bi-temporal
SIMILAR TO 0193fa2e-77c1 TOP 5            -- semantic: events nearest to this one
SEARCH 'risk' OF doc:* SIMILAR TO 'liability' ALPHA 0.5  -- hybrid

Aggregation & data shape

FACTS namespace, COUNT(*) OF * GROUP BY namespace
FACTS category, AVG(price_cents) OF sku:* GROUP BY category ORDER BY category
-- also SUM · MIN · MAX · MEDIAN · STDDEV · LISTAGG · COUNT(DISTINCT x)
-- · APPROX_COUNT_DISTINCT (HyperLogLog) · HAVING · LIMIT · OFFSET
PROFILE OF sku:*            -- one-shot data-shape summary of a subject set

Topology & data-shape analytics — a genuine differentiator

Operators that treat your facts as a shape, not just rows — borrowed from topological data analysis. No other database ships these:

SHAPE OF sku:* ON price_cents, margin   -- persistent-homology fingerprint (Betti numbers)
CONSISTENCY OF order:* EPS 0.1          -- sheaf-style cross-facet agreement
CYCLES IN CAUSES                          -- feedback loops in the causal graph
DRIFT OF price_cents BUCKETS 12          -- distribution drift over time

Point-in-time transactions

SNAPSHOT                               -- mark the current moment (a named chain head)
ROLLBACK OF sku:* TO 'yesterday'          -- append facts restoring prior state — history kept
DIFF OF sku:1 BETWEEN 'monday' AND NOW     -- exactly what changed between two moments

Writing (append-only)

PUT item:1 SET price_cents=799, note='spring sale'   -- supersedes the old fact
RETIRE item:1                                   -- a delete marker, history kept
DEFINE SCHEMA sku ...                            -- optional validation

AI inside the language

ENRICH asset:* USING vision           -- run a model over facts; the result is a fact
ASK 'which invoices from Acme are overdue?'  -- RAG answer with citations

Plain-English → CeQL isn't a statement you type — the ✨ Helper (and POST /v1/assist) translates English into a CeQL query you review before running: deterministic rules first, a local LLM for the long tail.

CePL — stored procedures

CePL procedures are stored as versioned facts and are self-tracing: every run records what it did, so the logic and its effects are auditable like everything else.

CeQL has an embedded, in-binary textbook (open /ceql on a running server) and an autocomplete catalog whose examples are tested to parse — the language documents and checks itself.

7 · SQL & the PostgreSQL wire protocol

For people and tools that speak SQL, Centauri offers a lean read-only SELECT that transpiles to CeQL — over REST (/v1/sql) and over the real PostgreSQL wire protocol:

# start the wire listener
centauri serve -pg-addr :5432

# now psql / JDBC / DBeaver / Tableau / Power BI connect directly
psql 'host=localhost port=5432 user=any sslmode=disable'
SELECT * FROM sku WHERE category='beverage' AS OF '2026-03-15' LIMIT 10;

Both the simple and the extended (prepared-statement) protocols are implemented, so JDBC and psycopg connect in their default mode. Honest scope: the wire path is read-only (writes use CeQL), every column comes back as text, and it serves the default database. It's a familiar front door for BI and LLMs, not a drop-in OLTP socket.

8 · Talking to AI & external SDKs

Centauri embeds no model. It talks to whatever model server you run, over the standard OpenAI-compatible HTTP shape — using only the Go standard library, so AI adds zero dependencies.

Models are facts

You register a model by writing a fact — so the choice is versioned, queryable, and auditable, not a hidden config file:

PUT model:vision FACET config SET
    endpoint='http://localhost:11434/v1/chat/completions',
    kind='vision', model='gemma3:27b', auth_env='OLLAMA_KEY'

kind is chat, embedding, vision, or image-embedding. auth_env names an environment variable holding the token — the secret is read at call time and never stored.

Works with

Local runtimes (recommended)

Ollama, LocalAI, vLLM — all OpenAI-compatible. centauri desktop (or serve -ai) auto-installs Ollama if missing and pulls a chat + embedder + vision trio sized to your hardware: small ~8 GB laptop (gemma3:4b + nomic-embed-text) · balanced 12–16 GB GPU (qwen3:14b + bge-m3) · max 24 GB+ (GLM-4.7-Flash, a 30B-MoE MIT-licensed GLM with 200K context, + bge-m3 + gemma3:27b vision).

Cloud models (opt-in)

Point an endpoint at any OpenAI-compatible API, or use the dashboard's one-click GLM-5.2 cloud boost (z.ai API key). Both are off by default and explicitly labeled, because the design is local-first — with a cloud model your questions leave the machine. GLM-5.2 cannot run locally: its weights need 200+ GB.

What the AI does over your data

  • RAG (ASK) — retrieves the most relevant facts (hybrid BM25 + vector) and has the model answer using only those, with citations to the exact source facts.
  • Auto-embed on ingest — with the appliance on, new facts embed themselves in the background, so everything is instantly searchable with no manual step.
  • Feedback loop — a thumbs-up/down on a source becomes a fact that re-ranks future retrieval. The system improves on your data with no model retraining.
  • Vision / extraction — read PDFs and images into structured, searchable facts.
  • Every inference is a fact — model, tokens, latency, result — so your AI's reasoning is part of the audit trail, not a black box.

Agents & programs

MCP server

Centauri ships a Model Context Protocol server over stdio, so AI agents are first-class clients: they query the same store directly for grounded memory and RAG — no glue code, no token cost just to remember.

SDKs

Zero-dependency clients for Python, Go, and JavaScript/TS (the Python SDK ships mock-server tests). REST is always there too, and Airbyte connectors move data in/out of 300+ systems.

The Genesis Engine — build a database from a conversation

Describe your scenario in plain language and Centauri interviews you and generates the whole blueprint (subjects, facets, schema), or paste existing SQL DDL to convert it. Then it stores that very conversation as facts — so the database forever remembers why it exists. Drive it from the dashboard (🏗 Build my database) or POST /v1/architect/plan & /v1/architect/apply.

9 · Enterprise & operations

NeedCentauri
Single sign-onOIDC/JWT verification (Okta, Azure AD, Auth0, Keycloak) — validates IdP tokens, with a write scope; stdlib only
Access controlToken auth, row-level security by subject prefix, and field-level masking (VPD-style)
High availabilityAutomatic failover via lease-based leader election with epoch fencing; /v1/ha reports role/leader
Retention & legal holdRETIRE-based retention policies; enforced legal hold blocks deletion of held subjects
ObservabilityPrometheus /metrics, /livez & /readyz probes, structured logs with request IDs, native TLS
Traffic protectionAdmission control: global and per-tenant concurrency caps, per-request timeouts, request-body caps, rate limiting
Cold storageS3-compatible tier (stdlib request signing), served on demand, Merkle-verified, LRU-cached
MultitenancyMany independent databases ("environments") per server, with snapshot cloning
Replication / CDCLog shipping + durable change-data-capture slots (resumable cursors)

10 · Scale & performance

  • Beyond RAM — a lazy disk-backed index keeps only the current fact per subject (plus small zone maps) resident; history, as-of, search and trace stream zone-map-pruned segments from disk, so total data can far exceed memory.
  • Write throughput — group commit coalesces fsyncs; sharding partitions subjects across independent logs (each its own chain + lock) and writes them in parallel.
  • Data skipping — zone-map statistics on each segment let the engine skip segments that can't match a predicate, without decompressing them.
  • Bounded recovery — periodic checkpoints + auto-seal keep open/recovery time flat regardless of total history; compaction merges small segments while preserving the chain.
  • Parallelism — independent segments and shards verify, compact and read in parallel across cores.
Honest limit: writes are single-writer per log and the hash chain is sequential by design — Centauri is not a high-contention multi-writer OLTP engine. Sharding scales throughput across subjects; it does not turn one chain into a parallel-write OLTP socket.

11 · How it compares

Centauri isn't trying to replace your operational database — it's the layer those engines don't have. It sits beside them as the system of record for what happened, when, why, how much to trust it, and now what the AI concluded.

Coming from Oracle? See the deep, concept-by-concept teardown — redo/undo/SCN/flashback/segments/RAC/Data Guard/TDE/VPD mapped to Centauri's internals, with side-by-side process walkthroughs: Centauri internals vs Oracle data management →

✓ built-in · ◐ partial / via add-ons · ✗ not its job. Comparisons are architectural; competitors evolve, so verify specifics for your version.
CapabilityCentauriPostgres
+ pgvector
MongoDBOraclePinecone /
Weaviate
Elastic-
search
Append-only, never-erased history✓✗✗◐✗✗
Bi-temporal (valid and belief time)✓✗✗◐✗✗
Causal lineage as first-class data✓✗✗✗✗✗
Topology / data-shape analytics (TDA)✓✗✗✗✗✗
Tamper-evidence (hash chain + Merkle)✓✗✗◐✗✗
Vector / semantic search✓✓◐◐✓◐
Ranked full-text (BM25)✓◐◐◐✗✓
Built-in local LLM / RAG with citations✓✗✗✗✗✗
SQL access◐ read✓✗✓✗✗
High-volume multi-writer OLTP✗✓✓✓✗✗
Agent-native (MCP)✓✗✗✗✗✗
Single binary, zero dependencies✓✗✗✗✗✗
Self-hosted, no license✓✓◐✗✗◐

vs Postgres + pgvector

They store vectors; you still wire the LLM, the audit trail, and the document pipeline yourself. Centauri ships all of it in one binary — and remembers why. Run it beside Postgres, not instead of it.

vs Oracle (Flashback / temporal tables)

Centauri is Oracle-grade in its lane — immutable, bi-temporal, tamper-evident, HA, SSO, SQL-wire — without the license, and without pretending to be a general OLTP engine. Offload OLTP to free Postgres, not a new license.

vs Pinecone / Weaviate

Their RAG lives in the cloud and keeps no audit trail. Centauri's vectors live next to bi-temporal, provenance-tagged facts, retrieval runs locally, and every answer cites replayable sources. For billion-scale pure ANN, a dedicated vector DB still wins.

vs Elasticsearch

Native BM25 + hybrid vector in one zero-dependency binary, bi-temporal, that explains why a hit ranked where it did. For massive corpora and advanced analyzers, a dedicated engine still wins.

12 · Where it fits — straight talk

✅ Reach for Centauri when

You need a system of record / audit trail, compliance-sensitive data, bi-temporal "as-of" reporting, or a private "ask your own documents" AI — and data ownership and a verifiable history matter.

↔ Pair it (don't force it) when

You need high-contention multi-writer OLTP or sub-millisecond transactions (use Postgres/SQLite — free — beside it), full read-write SQL over the wire, or billion-scale pure ANN.

Documentation honesty is policy: the comparison tables above mark what Centauri does not do, and we never inflate a ✗ to a ✓.

Your own AI · your own data · your own machine

See it for yourself

⬇ Download ★ GitHub Feature showcase Ask Centauri →