1 · The big picture
Centauri is one self-contained binary that is two things at once: an immutable system of record and a private AI engine that reasons over it — with zero third-party dependencies.
Ordinary databases store the latest value and overwrite the rest. Centauri stores every fact, forever, on an append-only log. An "update" appends a superseding fact; a "delete" appends a RETIRE marker. Because the past is never destroyed, Centauri can answer questions no ordinary store can — what was true at any moment, what you believed then, why it changed, and whether anyone tampered with it — and it keeps the same discipline for AI: every model answer and every human correction is itself a fact.
2 · The atom: a fact
The unit of storage is not a row or a document — it's an immutable fact (event). Each fact carries who/what it's about, the value, two clocks, and trust metadata:
- Subject — what the fact is about, e.g.
item:123456/store:4412. - Facet — which "reality" the fact belongs to (
source,register,shelf, …), so the same subject can hold several parallel truths. - Value — the payload (a flexible field map), optionally schema-checked.
- Two clocks — valid time (when it was true in the world) and transaction time (when Centauri learned it). This is what makes it bi-temporal.
- Provenance & confidence — where the fact came from and how much to trust it.
- Causal links — directed edges (
TRIGGERED,SUPERSEDES,CORRECTS,ENRICHED_FROM, …) so lineage is first-class data, not a join you invent later.
3 · The storage engine
A hot append-only log, sealed into compressed segments, tiered to cheap cold storage — with the working set indexed in RAM.
Write-then-apply
A write is appended to the log, linked into the hash chain, and made durable (fsync) before it changes any in-memory state. If the process crashes mid-write, recovery replays the log and lands in a consistent state — memory never gets ahead of disk. An optional group commit coalesces concurrent writes into a single fsync for throughput.
Hot log → sealed segments → cold tier
Recent facts live in a hot JSONL log. Once it grows past a threshold (or on demand) the tail is sealed into a segment that is flate-compressed (typically 5–10× on cold data) and carries uncompressed zone-map statistics + a Merkle root in its manifest — so the engine can prune and verify a segment without decompressing it. Sealed segments can be pushed to an S3-compatible cold tier (AWS/MinIO/R2/B2) over stdlib request signing; segments fetched from untrusted storage are Merkle-verified before use and LRU-cached.
Replay & determinism
Every in-RAM index (current-fact pointers, vectors, secondary indexes) is rebuilt by replaying the log through one apply step, so a restart reconstructs exactly the same state. A periodic checkpoint means open/recovery replays only the tail since the last checkpoint, not the whole history — recovery time stays bounded no matter how much data accumulates.
Crypto-erasure
For "right to be forgotten" without breaking the audit trail: a segment's payloads are encrypted with a per-segment AES-256-GCM key; destroying that key makes the payloads permanently unreadable while the hash chain — and therefore the proof of history — stays intact.
4 · Tamper-evidence (without a blockchain)
Every committed line is linked into a SHA-256 hash chain: each entry's hash folds in the previous entry's hash, so changing any past byte breaks every hash after it. Each sealed segment also publishes a Merkle root. The same chain is computed identically on the write path, on replay, and on raw ingestion (replication), and verify (or the dashboard's Verify button / /v1/verify) recomputes both to prove nothing was altered — with no downtime and no third party.
5 · Bi-temporal & causal queries
Two clocks unlock the questions ordinary databases can't answer:
| Question | How |
|---|---|
| What was true at a moment? | AS OF '2026-03-15' — valid-time travel |
| What did we believe at a moment? | AS KNOWN AT '2026-03-01' — transaction-time travel (the audit superpower) |
| Why did it change? | WHY <event> / EFFECTS <event> — walk the causal graph |
| Did a change actually land? | PENDING — find facts distributed but never activated |
| Where do sources disagree? | DISAGREE ON <field> — cross-facet conflict detection |
Causal links make lineage queryable directly: MATCH item:* CAUSES order:* finds chains across subjects without hand-written joins.
6 · The CeQL query language
CeQL is the native language where time, cause, and trust are syntax — not bolted-on functions. It comes in three interchangeable forms (text, JSON-AST, and REST) and has a plain-English translator for the rest.
Reading & time travel
FACTS OF item:* WHERE price_cents > 700 HISTORY OF item:1 -- the whole story of one subject FACTS OF item:1 AS OF YESTERDAY -- what was true then FACTS OF item:1 AS KNOWN AT '2026-03-01' -- what we believed then
Causality
WHY ev_01H... -- what caused this fact EFFECTS ev_01H... -- what this fact caused MATCH item:* CAUSES order:* -- causal chains across subjects
Search — keyword, vector, hybrid
SEARCH 'late markdown' OF item:* -- ranked BM25, bi-temporal SIMILAR TO 0193fa2e-77c1 TOP 5 -- semantic: events nearest to this one SEARCH 'risk' OF doc:* SIMILAR TO 'liability' ALPHA 0.5 -- hybrid
Aggregation & data shape
FACTS namespace, COUNT(*) OF * GROUP BY namespace FACTS category, AVG(price_cents) OF sku:* GROUP BY category ORDER BY category -- also SUM · MIN · MAX · MEDIAN · STDDEV · LISTAGG · COUNT(DISTINCT x) -- · APPROX_COUNT_DISTINCT (HyperLogLog) · HAVING · LIMIT · OFFSET PROFILE OF sku:* -- one-shot data-shape summary of a subject set
Topology & data-shape analytics — a genuine differentiator
Operators that treat your facts as a shape, not just rows — borrowed from topological data analysis. No other database ships these:
SHAPE OF sku:* ON price_cents, margin -- persistent-homology fingerprint (Betti numbers) CONSISTENCY OF order:* EPS 0.1 -- sheaf-style cross-facet agreement CYCLES IN CAUSES -- feedback loops in the causal graph DRIFT OF price_cents BUCKETS 12 -- distribution drift over time
Point-in-time transactions
SNAPSHOT -- mark the current moment (a named chain head) ROLLBACK OF sku:* TO 'yesterday' -- append facts restoring prior state — history kept DIFF OF sku:1 BETWEEN 'monday' AND NOW -- exactly what changed between two moments
Writing (append-only)
PUT item:1 SET price_cents=799, note='spring sale' -- supersedes the old fact RETIRE item:1 -- a delete marker, history kept DEFINE SCHEMA sku ... -- optional validation
AI inside the language
ENRICH asset:* USING vision -- run a model over facts; the result is a fact ASK 'which invoices from Acme are overdue?' -- RAG answer with citations
Plain-English → CeQL isn't a statement you type — the ✨ Helper (and POST /v1/assist) translates English into a CeQL query you review before running: deterministic rules first, a local LLM for the long tail.
CePL — stored procedures
CePL procedures are stored as versioned facts and are self-tracing: every run records what it did, so the logic and its effects are auditable like everything else.
/ceql on a running server) and an autocomplete catalog whose examples are tested to parse — the language documents and checks itself.7 · SQL & the PostgreSQL wire protocol
For people and tools that speak SQL, Centauri offers a lean read-only SELECT that transpiles to CeQL — over REST (/v1/sql) and over the real PostgreSQL wire protocol:
# start the wire listener centauri serve -pg-addr :5432 # now psql / JDBC / DBeaver / Tableau / Power BI connect directly psql 'host=localhost port=5432 user=any sslmode=disable' SELECT * FROM sku WHERE category='beverage' AS OF '2026-03-15' LIMIT 10;
Both the simple and the extended (prepared-statement) protocols are implemented, so JDBC and psycopg connect in their default mode. Honest scope: the wire path is read-only (writes use CeQL), every column comes back as text, and it serves the default database. It's a familiar front door for BI and LLMs, not a drop-in OLTP socket.
8 · Talking to AI & external SDKs
Centauri embeds no model. It talks to whatever model server you run, over the standard OpenAI-compatible HTTP shape — using only the Go standard library, so AI adds zero dependencies.
Models are facts
You register a model by writing a fact — so the choice is versioned, queryable, and auditable, not a hidden config file:
PUT model:vision FACET config SET endpoint='http://localhost:11434/v1/chat/completions', kind='vision', model='gemma3:27b', auth_env='OLLAMA_KEY'
kind is chat, embedding, vision, or image-embedding. auth_env names an environment variable holding the token — the secret is read at call time and never stored.
Works with
Local runtimes (recommended)
Ollama, LocalAI, vLLM — all OpenAI-compatible. centauri desktop (or serve -ai) auto-installs Ollama if missing and pulls a chat + embedder + vision trio sized to your hardware: small ~8 GB laptop (gemma3:4b + nomic-embed-text) · balanced 12–16 GB GPU (qwen3:14b + bge-m3) · max 24 GB+ (GLM-4.7-Flash, a 30B-MoE MIT-licensed GLM with 200K context, + bge-m3 + gemma3:27b vision).
Cloud models (opt-in)
Point an endpoint at any OpenAI-compatible API, or use the dashboard's one-click GLM-5.2 cloud boost (z.ai API key). Both are off by default and explicitly labeled, because the design is local-first — with a cloud model your questions leave the machine. GLM-5.2 cannot run locally: its weights need 200+ GB.
What the AI does over your data
- RAG (
ASK) — retrieves the most relevant facts (hybrid BM25 + vector) and has the model answer using only those, with citations to the exact source facts. - Auto-embed on ingest — with the appliance on, new facts embed themselves in the background, so everything is instantly searchable with no manual step.
- Feedback loop — a thumbs-up/down on a source becomes a fact that re-ranks future retrieval. The system improves on your data with no model retraining.
- Vision / extraction — read PDFs and images into structured, searchable facts.
- Every inference is a fact — model, tokens, latency, result — so your AI's reasoning is part of the audit trail, not a black box.
Agents & programs
MCP server
Centauri ships a Model Context Protocol server over stdio, so AI agents are first-class clients: they query the same store directly for grounded memory and RAG — no glue code, no token cost just to remember.
SDKs
Zero-dependency clients for Python, Go, and JavaScript/TS (the Python SDK ships mock-server tests). REST is always there too, and Airbyte connectors move data in/out of 300+ systems.
The Genesis Engine — build a database from a conversation
Describe your scenario in plain language and Centauri interviews you and generates the whole blueprint (subjects, facets, schema), or paste existing SQL DDL to convert it. Then it stores that very conversation as facts — so the database forever remembers why it exists. Drive it from the dashboard (🏗 Build my database) or POST /v1/architect/plan & /v1/architect/apply.
9 · Enterprise & operations
| Need | Centauri |
|---|---|
| Single sign-on | OIDC/JWT verification (Okta, Azure AD, Auth0, Keycloak) — validates IdP tokens, with a write scope; stdlib only |
| Access control | Token auth, row-level security by subject prefix, and field-level masking (VPD-style) |
| High availability | Automatic failover via lease-based leader election with epoch fencing; /v1/ha reports role/leader |
| Retention & legal hold | RETIRE-based retention policies; enforced legal hold blocks deletion of held subjects |
| Observability | Prometheus /metrics, /livez & /readyz probes, structured logs with request IDs, native TLS |
| Traffic protection | Admission control: global and per-tenant concurrency caps, per-request timeouts, request-body caps, rate limiting |
| Cold storage | S3-compatible tier (stdlib request signing), served on demand, Merkle-verified, LRU-cached |
| Multitenancy | Many independent databases ("environments") per server, with snapshot cloning |
| Replication / CDC | Log shipping + durable change-data-capture slots (resumable cursors) |
10 · Scale & performance
- Beyond RAM — a lazy disk-backed index keeps only the current fact per subject (plus small zone maps) resident; history, as-of, search and trace stream zone-map-pruned segments from disk, so total data can far exceed memory.
- Write throughput — group commit coalesces fsyncs; sharding partitions subjects across independent logs (each its own chain + lock) and writes them in parallel.
- Data skipping — zone-map statistics on each segment let the engine skip segments that can't match a predicate, without decompressing them.
- Bounded recovery — periodic checkpoints + auto-seal keep open/recovery time flat regardless of total history; compaction merges small segments while preserving the chain.
- Parallelism — independent segments and shards verify, compact and read in parallel across cores.
11 · How it compares
Centauri isn't trying to replace your operational database — it's the layer those engines don't have. It sits beside them as the system of record for what happened, when, why, how much to trust it, and now what the AI concluded.
Coming from Oracle? See the deep, concept-by-concept teardown — redo/undo/SCN/flashback/segments/RAC/Data Guard/TDE/VPD mapped to Centauri's internals, with side-by-side process walkthroughs: Centauri internals vs Oracle data management →
| Capability | Centauri | Postgres + pgvector | MongoDB | Oracle | Pinecone / Weaviate | Elastic- search |
|---|---|---|---|---|---|---|
| Append-only, never-erased history | ✓ | ✗ | ✗ | ◐ | ✗ | ✗ |
| Bi-temporal (valid and belief time) | ✓ | ✗ | ✗ | ◐ | ✗ | ✗ |
| Causal lineage as first-class data | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| Topology / data-shape analytics (TDA) | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| Tamper-evidence (hash chain + Merkle) | ✓ | ✗ | ✗ | ◐ | ✗ | ✗ |
| Vector / semantic search | ✓ | ✓ | ◐ | ◐ | ✓ | ◐ |
| Ranked full-text (BM25) | ✓ | ◐ | ◐ | ◐ | ✗ | ✓ |
| Built-in local LLM / RAG with citations | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| SQL access | ◐ read | ✓ | ✗ | ✓ | ✗ | ✗ |
| High-volume multi-writer OLTP | ✗ | ✓ | ✓ | ✓ | ✗ | ✗ |
| Agent-native (MCP) | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| Single binary, zero dependencies | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| Self-hosted, no license | ✓ | ✓ | ◐ | ✗ | ✗ | ◐ |
vs Postgres + pgvector
They store vectors; you still wire the LLM, the audit trail, and the document pipeline yourself. Centauri ships all of it in one binary — and remembers why. Run it beside Postgres, not instead of it.
vs Oracle (Flashback / temporal tables)
Centauri is Oracle-grade in its lane — immutable, bi-temporal, tamper-evident, HA, SSO, SQL-wire — without the license, and without pretending to be a general OLTP engine. Offload OLTP to free Postgres, not a new license.
vs Pinecone / Weaviate
Their RAG lives in the cloud and keeps no audit trail. Centauri's vectors live next to bi-temporal, provenance-tagged facts, retrieval runs locally, and every answer cites replayable sources. For billion-scale pure ANN, a dedicated vector DB still wins.
vs Elasticsearch
Native BM25 + hybrid vector in one zero-dependency binary, bi-temporal, that explains why a hit ranked where it did. For massive corpora and advanced analyzers, a dedicated engine still wins.
12 · Where it fits — straight talk
✅ Reach for Centauri when
You need a system of record / audit trail, compliance-sensitive data, bi-temporal "as-of" reporting, or a private "ask your own documents" AI — and data ownership and a verifiable history matter.
↔ Pair it (don't force it) when
You need high-contention multi-writer OLTP or sub-millisecond transactions (use Postgres/SQLite — free — beside it), full read-write SQL over the wire, or billion-scale pure ANN.
Documentation honesty is policy: the comparison tables above mark what Centauri does not do, and we never inflate a ✗ to a ✓.