Upgrade ArcadeDB

ArcadeDB automatically upgrades a database to a newer on-disk format when a newer version of ArcadeDB opens it. The migration is transparent and happens lazily on first open, so the only operator step is to point the new binaries at your existing data.

Breaking changes

This section lists behavior changes that may require action when upgrading.

26.11.1 - Client requests wait in a queue when the server is busy

Every request of a remote client (HTTP, Postgres, BOLT, Redis, gRPC, Gremlin Server, MCP, MongoDB) now goes through a query admission gate: at most twice the number of cores run at once, and a request that arrives when they are all taken, or when the running queries hold more than 80% of the query heap budget, waits in a queue and starts in arrival order (issue #9518). Before, every request started at once, and under a burst of heavy queries the last ones were refused by the heap budget or exhausted the heap.

Impact: under load a request can wait before it starts, up to arcadedb.queryQueueTimeout (30 seconds by default), and is refused as retryable once that runs out or when arcadedb.queryQueueMaxSize requests already wait (HTTP 503). A MongoDB request is refused at once instead of waiting. A BOLT result stream that is not read to the end keeps its slot until it is.

Action: none for most installations. Size the gate with arcadedb.queryMaxConcurrent, arcadedb.queryAdmissionHeapWatermark, arcadedb.queryQueueTimeout and arcadedb.queryQueueMaxSize (see Settings), or set arcadedb.queryMaxConcurrent=0 for the previous behavior.

26.11.1 - The Gremlin plugin fails the server start when its port is busy

If the Gremlin port (gremlin.port, default 8182) is already in use, the server now stops with an error. Before, it started without a Gremlin listener and still advertised the port in GET /api/v1/server (issue #9319). With gremlin.port=0 a free port is chosen at startup and advertised.

Impact: a server that used to start while another process held the Gremlin port no longer starts.

Action: free the port or set gremlin.port to a free one.

26.11.1 - A plain = / IN on a COLLATE CI index is case sensitive

The planner used to return every case variant for WHERE Name = 'JOHN' when a COLLATE CI index served the query, while the same predicate evaluated by a scan was case sensitive (issue #9403). The index no longer changes the answer: a plain = or IN is case sensitive with or without it.

Impact: a query that relied on Name = 'john' finding "John" through a COLLATE CI index now returns fewer rows.

Action: write the lookup on Name.toLowerCase(), for example WHERE Name.toLowerCase() = 'john', which the index serves. See Case-insensitive indexes.

26.11.1 - Cypher DURATION arithmetic is exact, and date * duration is an error

Durations stored in a property are read back as durations, including short ones such as P1D, so ORDER BY and comparisons on them now work (issue #9338). Duration arithmetic raises an arithmetic error when the result does not fit, instead of returning a wrong value. A date or time combined with a duration using *, / or % is now a type error, and duration + date is accepted.

Impact: a query that relied on a stored duration being plain text, or on date * duration returning a date, changes.

Action: none for normal use. A range index on a duration property still orders by its text form (issue #9348).

26.11.1 - Sparse-vector top-K is split across workers from 50,000 postings

vector.sparseNeighbors on an LSM_SPARSE_VECTOR index now splits a query across the scoring workers once its terms hold 50,000 postings in total (setting arcadedb.sparseVectorScoringMinPostingsForPartitioning, default 50000, was 200000). Queries between 50,000 and 200,000 postings used to run on one thread and set the slowest latencies. A query just past the threshold uses two ranges and leaves the other workers free; wider queries use more (issue #9482).

Impact: results are the same. Queries in that size range are faster, and use more CPU when the server is otherwise idle.

Action: none. On a small or shared server where CPU matters more than latency, raise arcadedb.sparseVectorScoringMinPostingsForPartitioning (for example back to 200000).

26.11.1 - Production mode conceals error text on Bolt, Redis, /ws and MCP

In production server mode the /ws insert session, /ws change-event subscriptions, Bolt, Redis and MCP now answer an engine failure with "The request failed. Check the server log for the details" and write the detail to the server log, as HTTP, gRPC, PostgreSQL, MongoDB and Gremlin Server already did (issue #8749). Before, they returned the engine’s message as is, including the stored values of a duplicated key.

Impact: a client of these protocols that reads the error text in production gets the generic message. Error codes, Neo4j status codes, RESP prefixes and exception class names are unchanged, and so is text the protocol writes about the request itself.

Action: branch on the code rather than on the message, and read the detail from the server log. See Error details in production mode.

26.11.1 - Graph algorithms answer the incoming side of unidirectional edge types

The algo. and path. procedures, the path-finding SQL functions (dijkstra(), astar(), bellmanFord(), cchShortestPath(), duanSSSP()) and the node.degree* / node.relationship.* functions now follow the edges of a UNIDIRECTIONAL type from their target too, as a Graph Analytical View already did (issue #8629). Before, they saw such an edge from its source only, so for example a walk with direction IN or BOTH missed it. refactor.mergeNodes, refactor.cloneNodesWithRelationships and changing a node’s labels in Cypher no longer drop the incoming unidirectional edges of the node.

Impact: results over unidirectional edge types can change (more paths, higher in-degrees, different scores), and now match the results with a Graph Analytical View.

Action: none. See CREATE TYPE for the heap a walk from the target takes.

26.11.1 - Cypher counts match an undirected self loop once

An undirected relationship pattern such as (n)-[:KNOWS]-(m) matches a self loop (an edge from a vertex to itself) once, as the openCypher specification and Neo4j do. The rows of such a pattern always did, but the counts computed without materializing the rows counted it twice: MATCH (a)-[:KNOWS]-(b) RETURN count(), a count() grouped by the start node, OPTIONAL MATCH …​ WITH n, count(m) and the COUNT { } subquery, with and without a Graph Analytical View (issues #8750, #9540). The same release corrects the count of a chain with an inequality such as WHERE a <> c when the chain has a labelled or multi-hop part before or after the two positions the inequality compares.

Impact: those counts can be lower than before on graphs with self loops, or on such chains, and now equal the number of rows the same MATCH returns.

Action: none.

26.11.1 - Cypher count(*) of an OPTIONAL MATCH with several arms, and of a Cartesian product

A count(*) over a star pattern no longer counts an OPTIONAL MATCH that reaches the central node from two sides, or holds two patterns, as two independent optional matches: when one side finds nothing the clause contributes one row, not the number of matches of the other side. A star next to a node pattern that shares no variable with it, MATCH (a)-[:R]→(b), (a)-[:S]→(c), (t:Tag), now multiplies by the Tag count, which it used to leave out. Both were answered by the count push-down without building the rows (issue #9596).

count(*) over parts of a MATCH that share no variable is now the product of the parts' counts and builds no row, so a count whose result does not fit a 64-bit integer raises an error instead of running until the command timeout.

Impact: the counts of those two star shapes change, and now equal the number of rows the same MATCH returns.

Action: none.

26.10.1 — Sparse-vector top-K scores a window of documents at a time

vector.sparseNeighbors on an LSM_SPARSE_VECTOR index now scores documents in windows (setting arcadedb.sparseVectorScoringWindow, default 16384). On wide queries this is about twice as fast, and queries no longer wait for a COMPACT INDEX or a memtable flush to finish (issues #9200, #9209, #9210).

Impact: scores can differ from earlier versions in the last bit, so documents with exactly equal scores may rank in a different order.

Action: none. Set arcadedb.sparseVectorScoringWindow to 0 to go back to the previous traversal.

26.10.1 — The remote ArcadeGraph connects to the Gremlin port the server is bound to

The remote ArcadeGraph client used to connect to Gremlin Server on the fixed port 8182. It now uses the port the server advertises (GET /api/v1/server?mode=cluster, ports.gremlin), and falls back to 8182 when the server advertises none (issue #8578).

Impact: only a client that reaches Gremlin Server through a port mapping (for example a container or load balancer publishing 8182 to a different internal port), which worked by accident before.

Action: set the client setting arcadedb.gremlin.client.port to the port the client must connect to.

26.10.1 — Negative zero equals zero, and a few query results change

-0.0 is now equal to 0.0 everywhere (SQL and Cypher, =, IN, ordering, indexed or not), as IEEE 754 and openCypher define it.

Action: a HASH index over a FLOAT or DOUBLE property that already stores -0.0, or an LSM index that stored both -0.0 and 0.0, should be rebuilt once after upgrading:

REBUILD INDEX `Measure[value]`

Other results that change:

  • vector.neighbors, vector.sparseNeighbors and db.index.vector.queryNodes on a parent type spec (Parent[prop]) now also search the records of its sub-types. Use the sub-type name in the spec to search only that sub-type.

  • A FLOAT property is equal to the DOUBLE bound that narrows to it, so <, ⇐, > and >= agree with = and with an index (v = 16777217.0 finds 16777216f, and v < 16777217.0 no longer does).

  • An integer bound above 2^53 on a DOUBLE property, or above 2^24 on a FLOAT property, is answered by a scan instead of the index, so the result is exact.

26.10.1 — HASH indexes on a DECIMAL property need one REBUILD INDEX after upgrading

A DECIMAL keeps the scale it was written with, so 5 and 5.00 are the same number stored two different ways. The LSM_TREE index family already treated them as one key; the HASH family did not, because it settles key identity on the serialized bytes and the two spell out differently. A UNIQUE_HASH index on a DECIMAL property therefore accepted both as separate rows, and a lookup for one spelling could not find a row stored under the other. Fixed in 26.10.1 (issue #7767).

Impact: any HASH/UNIQUE_HASH/NOTUNIQUE_HASH index on a DECIMAL property created before 26.10.1. LSM_TREE/UNIQUE/NOTUNIQUE indexes are unaffected.

Action: run REBUILD INDEX once, after upgrading, on every HASH index over a DECIMAL property:

REBUILD INDEX `Invoice[amount]`

Until it is rebuilt, the index file still holds the keys in their original form and lookups stay unreliable. If duplicates slipped in while a UNIQUE_HASH constraint was not being enforced, the rebuild reports them as a duplicated key — correct the data, then rebuild again.

26.10.1 — Case-insensitive HASH indexes need one REBUILD INDEX after upgrading

A HASH/UNIQUE_HASH/NOTUNIQUE_HASH index created with COLLATE CI did not actually fold case: an existing key and its differently-cased duplicate could both be inserted into a UNIQUE_HASH index, and a case-insensitive query answered through such an index could return fewer rows than the same query without an index. Fixed in 26.10.1 (issue #7766).

Impact: any HASH-family index with COLLATE CI created before 26.10.1. LSM_TREE/UNIQUE/NOTUNIQUE indexes with COLLATE CI already folded case correctly and are unaffected.

Action: run REBUILD INDEX once, after upgrading, on every case-insensitive HASH index:

REBUILD INDEX `Product[name]`

This rewrites the index with the keys folded correctly. See Case-insensitive indexes.

26.10.1 - Writes that used to store NULL or a coerced value now fail

A value that the declared property type cannot take is refused with an error naming the property, instead of being stored silently as NULL or coerced (issues #8090, #9110). This covers:

  • an unparsable datetime string written to a DATETIME/DATE property (the SQL timestamp form 2024-02-29 13:45:10.123456, which PostgreSQL clients send, is now parsed instead of being stored as NULL);

  • a number that overflows a LONG, FLOAT or DOUBLE property, an empty string written to a number, a fraction written to a BOOLEAN, and wrong-typed values written to DECIMAL, DATETIME, BINARY or LINK properties;

  • a typed property set from a bare sub-query, for example UPDATE Order SET id = (SELECT max(id) FROM Order), which used to store NULL and now fails with Cannot convert type ArrayList to INTEGER.

Schemaless properties are unaffected.

Action: fix the statements that now fail. For a sub-query, extract the scalar, for example SET id = (SELECT max(id) AS m FROM Order)[0].m.

26.10.1 - A write computed from a stale read is refused

A property write on a record that another transaction changed after the current transaction read it is refused with a retryable ConcurrentModificationException, instead of silently overwriting the concurrent change (issue #8610). SQL UPDATE and Cypher SET/MERGE/REMOVE are not affected, since they compute their values from the latest data. See Isolation.

Action: retry the transaction on ConcurrentModificationException (the database.transaction(…​, retries) API does). Set arcadedb.txStaleReadCheck=false to restore the previous last-writer-wins behavior.

26.10.1 - TRUNCATE TYPE refuses a vertex or edge type that holds records

The refusal was documented but never fired, so TRUNCATE TYPE silently emptied live vertex and edge types (issue #8042).

Action: add UNSAFE where the truncation is intended. See TRUNCATE TYPE.

26.10.1 - Postgres wire: ROLLBACK TO SAVEPOINT is refused

It used to answer success while keeping every write made since the savepoint, so the next COMMIT persisted them. It now fails with SQLSTATE 0A000 and the transaction must be rolled back. See Savepoints.

26.10.1 - Stricter authorization for schema DDL and Gremlin scripting

  • ALTER TYPE (including the edge-type lightweight and unique settings) and REBUILD TYPE require the updateSchema permission (security advisory GHSA-9c7g-grf7-j2r5).

  • Gremlin capabilities that run code or touch files on the host are reserved to root: the groovy engine (and the auto fallback to it) over HTTP, Gremlin Server scripts in a language other than gremlin-lang, bytecode that carries a lambda, and the io() step on every engine (security advisory GHSA-h2j4-28h8-cj5v). Plain gremlin-lang traversals are unchanged. Embedded use without a bound user is unaffected.

Action: grant updateSchema to the users that manage types, and run Groovy scripts, lambdas and io() as root or rewrite them as gremlin-lang traversals. See Gremlin security.

26.10.1 - API tokens minted over plain HTTP are logged

POST /api/v1/server/api-tokens over plain HTTP from a non-loopback peer still works, but logs a WARNING. The new gRPC CreateApiToken RPC always requires TLS or a loopback peer. With arcadedb.server.apiTokenRequireSecureTransport=true it is refused with 412, and the default of that setting is scheduled to change to true in 27.1.1.

Action: mint tokens over HTTPS, from localhost, or through a reverse proxy listed in arcadedb.server.apiTokenTrustedProxies. See Create an API token and, for the Docker first-run flow on http://localhost:2480, Docker.

26.10.1 - HA defaults that changed

Setting Before 26.10.1 From 26.10.1

ha.securityConvergenceReadinessTimeout

no gate

With server.readinessRequiresHA on, a node that joined at runtime or installed a snapshot reports NOT READY for up to 30 s until its users, groups and API tokens match the leader’s (issues #7819, #8414). 0 disables it.

ha.securityEntryCapabilityGate

n/a

true: creating or changing a group or API token is refused with 409 while any peer is down or still runs an older version, so a rolling upgrade never halts an old node. It succeeds once every node is up on 26.10.1.

ha.txSchemaCheck

n/a

true: a replica transaction prepared before a schema change the leader already applied is refused with a retryable ConcurrentModificationException. Expect retries on replicas during DDL.

ha.divergedFollowerRecoveryDurationMs

shared ha.staleFollowerRecoveryDurationMs (60000)

separate window, default 20000. A deployment that raised the old setting to delay the reformat of a diverged follower must raise this one.

ha.grpcMaxConnectionIdleMs

unbounded

300000: an idle inbound Raft gRPC connection is closed with a GOAWAY and reopened on the next call. 0 restores the old behavior.

ha.jvmPauseCloseThresholdMs

Ratis closed the Raft division after a 60 s JVM pause

0: the division is not closed; a long pause still steps a leader down and the health monitor decides on recovery. Set 60000 to restore the old behavior.

See HA cluster upgrade for the rolling-upgrade procedure.

26.9.1 — Sparse-vector indexes need one REBUILD INDEX after upgrading

A partial compaction of an LSM_SPARSE_VECTOR index could give the merged segment a position ahead of a segment it had not merged. When that skipped segment held the tombstone for a document an older merged segment still had a live posting for, the delete was overruled and the document reappeared in vector.sparseNeighbors results — permanently, because the next full compaction then discarded the tombstone as the loser. The same inversion could restore a stale weight after an update. Fixed in 26.9.1 (issue #6379).

Impact: LSM_SPARSE_VECTOR indexes on databases created before 26.9.1 that have seen deletes or updates. Upgrading corrects the ordering for every new merge, but a segment already written under the old rule cannot be repaired automatically — the record of where it should have sat is gone.

How to tell: the server reports it once per affected index, when the index is opened:

Sparse-vector index 'Doc[tokens,weights]' holds 3 segment(s) whose precedence predates the
recency-epoch fix (issue #6379); a document deleted before those segments were merged may still
be returned by queries. Run 'REBUILD INDEX Doc[tokens,weights]' once to rewrite them in the
correct order.

The flag behind that message is stored in the segment and inherited by any later merge, so ordinary write traffic compacting the affected segments away does not silence it. Nothing else distinguishes an affected index: queries succeed and the data looks intact.

Action: run once per sparse-vector index, after upgrading:

REBUILD INDEX `Doc[tokens,weights]`

This collapses the existing segments into a single correctly ordered segment. Indexes that have only ever been appended to are unaffected, as are databases created on 26.9.1 or later. See Sparse Vector Search.

HA rolling upgrade: the corrected ordering is recorded in the segment file, but reading it correctly requires the fix. During a staged rollout, a node still on the older build that receives a segment merged by an upgraded leader falls back to ordering by segment id and can still return a deleted document, until it is upgraded too. Upgrade every node before running the REBUILD INDEX above, and treat sparse-vector query results from a mixed-version cluster as provisional.

Single-server upgrade

  1. Download and extract the new ArcadeDB release in a separate directory (do not overwrite the old install).

  2. Stop the running server (or close every open database via the HTTP close database command).

  3. Copy the directories listed in What to copy from the old install into the new install, preserving the original paths.

  4. Start the new server.

When the server starts, every database is opened with the new code and any required format migration is applied in place.

What to copy

ArcadeDB’s root directory contains a fixed set of folders. Only a few of them hold user-owned state; the rest are recreated from the release archive on every install.

Folder Action Notes

databases/

Copy

The actual database data. Every subdirectory is one database. This is the only folder strictly required for an upgrade.

config/

Copy selected files

The release ships a default config/ directory. Carry over only the files you have actually modified — listing below.

backups/

Copy if you want history

Backup archives produced by manual or automatic backups. The server does not need them to start; copy them only if you want to keep the backup history reachable from the Studio Backup tab and from RESTORE DATABASE commands.

log/

Skip

Server log files. Safe to start with an empty log/ directory; the release ships an empty one. Copy only if you want to keep historical logs alongside the new install.

raft-storage-<peerId>/ (HA only)

Copy

Per-node Raft log segments, created at runtime when HA is enabled. This storage is durable by default (arcadedb.ha.raftPersistStorage=true) and must be preserved so a node can rejoin by replaying its log instead of a full snapshot resync. When raftPersistStorage is explicitly false, it is ephemeral and can be skipped. See HA cluster upgrade for details.

replication/

Skip

Legacy directory from the pre-Raft replication implementation. Not used by the current HA stack.

lib/, bin/

Use the new release as-is

Java artifacts and start scripts. Always take them from the new release — do not copy from the old install. The only exception is custom JARs you have dropped into lib/ yourself; carry those over by hand.

Files to carry over from config/

The config/ directory ships with a small set of defaults. You only need to copy the files that you have actually changed:

File When to copy

server-users.jsonl

Always, unless you create users only via REST/SQL on every upgrade. This is where ArcadeDB persists local users and their per-database group assignments.

server-groups.json

If you have customised the security policy (custom groups, type-level permissions, result-set limits). See Security Policy.

backup.json

If you have configured the auto-backup scheduler (schedule, retention, target directory). See Automatic Backup.

arcadedb-log.properties

If you have tuned java.util.logging settings (log levels, rolling file size, etc.).

gremlin-server.yaml, gremlin-server.groovy, gremlin-server.properties

Only if you use the embedded Gremlin Server and have edited any of these.

Everything else under config/ (arcadedb-log-all.properties, arcadedb-statefulset.yaml, …) ships unchanged with each release; take the new version.

ArcadeDB does not read a config/server-configuration.json file in normal operation. Server-wide settings come from JVM system properties, environment variables and the arcadedb-server.sh arguments — see Server Configuration.

HA cluster upgrade

An HA cluster (see High Availability) is upgraded one node at a time. The per-node procedure is the same as single-server above, with the following clarifications.

What to migrate on each node:

  • databases/ — copy, exactly as in the single-server case.

  • config/ — copy the same files as in the single-server case. The cluster’s identity is derived from arcadedb.ha.clusterName and the root password (the inter-node token is computed from them at startup), so as long as every node keeps the same cluster name and the same root password, the upgraded cluster reattaches to itself without any extra step.

  • raft-storage-<peerId>/ — copy. The local Raft log is durable by default (arcadedb.ha.raftPersistStorage=true), so preserve this directory across the upgrade to let the node rejoin by replaying its own log. Whether or not it is preserved, after the upgraded node rejoins the cluster the leader replays any missing entries via Raft or, if the node has fallen behind the purge boundary, streams a full snapshot over HTTP automatically. With raftPersistStorage=false this directory is ephemeral and can be skipped.

  • backups/ and log/ — same as single-server.

Recommended rolling upgrade flow (zero downtime, requires quorum at all times):

  1. Pick a non-leader node first.

  2. Stop the node, install the new release in a separate directory, copy the folders listed above, start the new binary.

  3. Wait until GET /api/v1/cluster reports the node as healthy and caught up (replication lag near zero).

    1. The Studio Cluster tab shows the same information visually.

  4. Repeat for every remaining replica.

  5. Step the leader down with POST /api/v1/cluster/stepdown (or simply stop it — Raft will elect a new leader) and upgrade it last.

From 26.10.1, a node that receives replicated TIMESERIES data for a type it does not have — or for a type that exists locally but is not a TIMESERIES one — refuses the entry, which quarantines that database and triggers an automatic resync from the leader. Earlier releases logged the problem and carried on, which left the node’s copy of that type permanently behind the leader’s with nothing but a log line to say so.

This should not occur during an ordinary rolling upgrade, since the entry that creates a type is always applied before data for it arrives. It is a safeguard for the case where it does — for instance a mixed-version window in which one node’s build cannot construct a type another node has. If you see a database quarantined with a message about a TimeSeries sealed store, the resync will restore it; complete the rollout so every node runs the same build.

For first-time HA bring-up with a pre-existing database (e.g. when you scale a single-node install into a cluster), see Offline Cluster Bootstrap.

Docker / Kubernetes

When you run ArcadeDB from the official container image, the release is replaced atomically by changing the image tag. The only data that survives across container restarts is what lives on a persistent volume, so:

  • Mount /home/arcadedb/databases on a persistent volume — this is the equivalent of "copy the databases/ folder".

  • Mount /home/arcadedb/config only if you have edited the files listed in Files to carry over from config/; otherwise let the image ship its defaults.

  • Set arcadedb.ha.raftStorageDirectory to a static path (e.g. /home/arcadedb/raft-storage) and mount that path on a persistent volume. The default per-node directory is named raft-storage-<nodeName> dynamically, which a container runtime or Kubernetes cannot mount as a volume (a wildcard like /home/arcadedb/raft-storage-* is created literally, not expanded) — this is exactly what raftStorageDirectory is for. The Raft log is durable by default (arcadedb.ha.raftPersistStorage=true) so a restarted pod can replay its own log instead of resyncing from peers. With raftPersistStorage=false, the log is ephemeral and this mount is unnecessary.

  • Mount /home/arcadedb/backups if you want to retain backup archives across pod restarts.

For Kubernetes specifics (StatefulSet, headless service, init container for pre-staging a database), see Kubernetes.

Downgrade ArcadeDB

In case you need to downgrade to an older version of ArcadeDB, check the binary compatibility between the versions. ArcadeDB uses semantic versioning with 100% on-disk compatibility for migration of databases up or down between patch versions (the Z in X.Y.Z). For minor or major downgrades the safest path is to export the database with the newer version and re-import it with the older version.

Downgrading a node of an HA cluster

A cluster running v26.10.1 or later replicates schema changes in a compact format that earlier builds cannot read. Those entries stay in the Raft log, so a node that is restarted on an older build replays them and ends up with a schema that no longer matches the rest of the cluster - silently.

A node of a cluster that has run v26.10.1 or later must not be rolled back below it by restarting it on its existing Raft log. If you have to go back that far, rebuild the node instead: stop it, delete its raft-storage-<peerId>/ directory and its databases/ folder, and start the older build so it rejoins empty and takes a fresh copy from the leader.

This applies only to the Raft log, not to your data. Database files themselves keep the usual patch-version compatibility described above.