Storage Internals
Page Version
Records are stored in pages. Each page has its own version number, which increments on each update. At creation the page version is zero. In optimistic transactions, ArcadeDB checks the version in order to avoid conflicts at commit time.
Schema File Persistence
The schema of a database lives in two JSON files in the database directory: schema.json, the current generation, and schema.prev.json, the generation it replaced.
Every DDL statement rewrites them, once (see Cost of a Schema Change, and Batching Many).
Since v26.10.1, both files are published by an atomic replacement: the new content is written to a temporary sibling file, flushed to disk, and then renamed over its target.
A process reading schema.json from the filesystem - an external backup, a monitoring script, a snapshot tool - therefore always observes one complete generation.
It is never truncated, and it never disappears.
Before v26.10.1, saving the schema renamed schema.json aside to schema.prev.json and only then wrote the replacement.
For the duration of that write, schema.json did not exist, and a process killed inside that window left a database whose schema file was missing and recoverable only from schema.prev.json.
If you have tooling that copies a database directory while the database is running, and that tooling tolerates a missing or empty schema.json, that workaround is no longer needed.
|
Since v26.10.1 a full backup archives schema.prev.json as well, so a restored database keeps the same fallback copy as the database it came from.
A database whose schema has never been re-saved has no previous generation yet, and the archive simply does not contain one.
The previous generation is saved by creating a second hard link to the bytes already on disk where the file system supports it, so schema.json and schema.prev.json can share one inode until the next DDL statement replaces schema.json.
This is invisible to anything that reads or copies the files - cp, tar, a backup archive and a restore all produce two independent files - but a tool that reports inode identity or disk usage will see the two names counted once.
File systems without hard links (FAT/exFAT, some network mounts) fall back to a byte copy automatically.
On a file system that cannot rename atomically at all, ArcadeDB falls back to a plain replacement and logs a warning once per JVM. The schema still saves; only the crash-time guarantee is lost.
Cost of a Schema Change, and Batching Many
Each schema write is durable: the new schema.json is flushed to disk and the directory that holds it is flushed once, after both files are published.
On a real disk that is a few milliseconds to a few tens of milliseconds per write, depending on the device, against a fraction of a millisecond on tmpfs.
Since v26.10.1 every DDL statement writes the schema exactly once, however many bucket and index files and internal steps it involves.
Before, a CREATE TYPE wrote it five or six times, a CREATE PROPERTY or a CREATE INDEX twice, and each write flushed the directory twice.
An application or a test suite that creates hundreds of types pays that write once per statement. To pay it once for a whole batch:
-
SQL: send the statements as one
sqlscript. A script made only of schema definition statements runs as one schema session and writes the schema once (settingschemaBulkDDLScript, on by default). Under HA the batch is also replicated as one Raft entry. -
Java: wrap the calls in
db.getSchema().bulkChange(() → { … }), which does the same for any sequence of schema API calls. -
Any API: run the DDL inside a transaction. The schema write is postponed to the end of the transaction.
Creating the bucket and index files themselves is cheap: about a millisecond per type with one index, once the schema writes are batched.
Schema changes are not transactional.
A DDL statement run inside a transaction takes effect immediately and is visible to every other session; a rollback does not undo it.
Its types, properties and files stay, and the schema is written at the end of the transaction whether it commits or rolls back.
A crash before that write leaves the new bucket and index files on disk without a schema entry naming them: the database opens normally without the new types, and running the same DDL again succeeds.
The window lasts as long as the transaction, so it is worth keeping a transaction that creates types or indexes short, and creating them before a long import rather than inside it.
Records another session committed into such a type before the crash are in those files too, and are not reachable after the restart: recreate the type and they stay unreferenced.
Because a schema change is not transactional, the next commit on any session writes it, including a session that only writes records, and the type it names may belong to a transaction that has not ended yet.
The transaction that ran the DDL is not exposed to this for its own records: the schema is written just before its commit becomes durable, so begin(); CREATE TYPE; INSERT; commit() never leaves acknowledged records in files the schema does not name.
Under HA, each DDL statement inside a transaction is still replicated on its own, so prefer a script or bulkChange() there.
|