Transactions
A transaction comprises a unit of work performed within a database management system (or similar system) against a database, and treated in a coherent and reliable way independent of other transactions. Transactions in a database environment have two main purposes:
-
to provide reliable units of work that allow correct recovery from failures and keep a database consistent even in cases of system failure, when execution stops (completely or partially) and many operations upon a database remain uncompleted, with unclear status
-
to provide isolation between programs accessing a database concurrently. If this isolation is not provided, the program’s outcome are possibly erroneous.
A database transaction, by definition, must be atomic, consistent, isolated and durable. Database practitioners often refer to these properties of database transactions using the acronym ACID). - Wikipedia
ArcadeDB is an ACID compliant DBMS.
| ArcadeDB keeps the transaction in the host’s RAM, so the transaction size is affected by the available RAM (Heap memory) on JVM. For transactions involving many records, consider to split it in multiple transactions. |
Atomicity
"Atomicity requires that each transaction is 'all or nothing': if one part of the transaction fails, the entire transaction fails, and the database state is left unchanged. An atomic system must guarantee atomicity in each and every situation, including power failures, errors, and crashes. To the outside world, a committed transaction appears (by its effects on the database) to be indivisible ("atomic"), and an aborted transaction does not happen." - Wikipedia
Consistency
"The consistency property ensures that any transaction will bring the database from one valid state to another. Any data written to the database must be valid according to all defined rules, including but not limited to constraints, cascades, triggers, and any combination thereof. This does not guarantee correctness of the transaction in all ways the application programmer might have wanted (that is the responsibility of application-level code) but merely that any programming errors do not violate any defined rules." - Wikipedia
ArcadeDB uses the MVCC to assure consistency by versioning the page where the record are stored.
Look at this example:
| Sequence | Client/Thread 1 | Client/Thread 2 | Version of page containing record X |
|---|---|---|---|
1 |
Begin of Transaction |
||
2 |
read(x) |
10 |
|
3 |
Begin of Transaction |
||
4 |
read(x) |
10 |
|
5 |
write(x) |
10 |
|
6 |
commit |
10 → 11 |
|
7 |
write(x) |
10 |
|
8 |
commit |
10 → 11 = Error, in database x already is at 11 |
Isolation
"The isolation property ensures that the concurrent execution of transactions results in a system state that would be obtained if transactions were executed serially, i.e. one after the other. Providing isolation is the main goal of concurrency control. Depending on concurrency control method, the effects of an incomplete transaction might not even be visible to another transaction." - Wikipedia
The SQL standard defines the following phenomena which are prohibited at various levels are:
-
Dirty Read: a transaction reads data written by a concurrent uncommitted transaction. This is never possible with ArcadeDB.
-
Non Repeatable Read: a transaction re-reads data it has previously read and finds that data has been modified by another transaction (that committed since the initial read).
-
Phantom Read: a transaction re-executes a query returning a set of rows that satisfy a search condition and finds that the set of rows satisfying the condition has changed due to another recently-committed transaction. This happens also when records are deleted or inserted during the transaction and they could become visible during the transaction.
The SQL standard transaction isolation levels are described in the table below:
| Isolation Level | Dirty Read | Non repeatable Read | Phantom Read |
|---|---|---|---|
|
Not possible |
Possible |
Possible |
|
Not possible |
Not possible |
Possible |
The SQL SERIALIZABLE level is not supported by ArcadeDB.
Under REPEATABLE_READ the pages a transaction reads are kept and read again from there, so a record read twice returns
the same content. A record larger than a page spans several pages, and the transaction may already hold its first page
from reading another record stored there. If another transaction rewrites the large record before it is read, the version
that page belongs to can no longer be read whole, and the read fails with a retryable ConcurrentModificationException
rather than returning parts of two versions. Retrying the transaction reads the new version (since v26.11.1).
Lost updates. Under READ_COMMITTED, reading a record and then writing a value computed from it is protected against a
concurrent change: if another transaction committed a change to the record after you read it, saving a property on it
fails with a retryable ConcurrentModificationException, instead of silently overwriting that change. Retrying the
transaction, for example with database.transaction(…), reads the record again and applies the write on the new value.
This holds for records of any size, including documents larger than a page. Likewise, a DELETE that selected a record which
another transaction then deleted and committed fails with a retryable ConcurrentModificationException, so the retry simply
no longer finds the record.
database.transaction(() -> {
Vertex v = rid.asVertex();
v.modify().set("n", v.getInteger("n") + 1).save(); // retried, never lost
});
Adding or removing edges on a vertex is never refused, and a record kept from an earlier transaction is refreshed as
before. The check can be disabled with the txStaleReadCheck setting (since v26.10.1).
Edges to a deleted vertex. Creating an edge of a UNIDIRECTIONAL type to a vertex that another transaction deleted and
committed meanwhile fails with a retryable ConcurrentModificationException, as it already did for bidirectional edges, so
no edge is left pointing at a deleted vertex. Likewise, a transaction deleting a vertex fails and can be retried if another
transaction committed an edge of a UNIDIRECTIONAL type ending in that bucket after the delete looked for its incoming
edges (since v26.10.1).
Using remote access all the commands are executed on the server, so out of transaction scope.
Look below for more information.
Look at these examples:
| Sequence | Client/Thread 1 | Client/Thread 2 |
|---|---|---|
1 |
Begin of Transaction |
|
2 |
read(x) |
|
3 |
Begin of Transaction |
|
4 |
read(x) |
|
5 |
write(x) |
|
6 |
commit |
|
7 |
read(x) |
|
8 |
commit |
At operation 7 the client 1 continues to read the same version of x read in operation 2.
| Sequence | Client/Thread 1 | Client/Thread 2 |
|---|---|---|
1 |
Begin of Transaction |
|
2 |
read(x) |
|
3 |
Begin of Transaction |
|
4 |
read(y) |
|
5 |
write(y) |
|
6 |
commit |
|
7 |
read(y) |
|
8 |
commit |
At operation 7 the client 1 reads the version of y which was written at operation 6 by client 2. This is because it never reads y before.
Durability
"Durability means that once a transaction has been committed, it will remain so, even in the event of power loss, crashes, or errors. In a relational database, for instance, once a group of SQL statements execute, the results need to be stored permanently (even if the database crashes immediately thereafter). To defend against power loss, transactions (or their effects) must be recorded in a non-volatile memory." - Wikipedia
Fail-over
An ArcadeDB instance can fail for several reasons:
-
Hardware problems, such as loss of power or disk error
-
Software problems, such as a operating system crash
-
Application problems, such as a bug that crashes your application that is connected to the ArcadeDB engine.
You can use the ArcadeDB engine directly in the same process of your application. This gives superior performance due to the lack of inter-process communication. In this case, should your application crash (for any reason), the ArcadeDB engine also crashes.
If you’re connected to an ArcadeDB server remotely, and if your application crashes but the engine continues to work, any pending transaction owned by the client will be rolled back.
Auto-recovery
At start-up the ArcadeDB engine checks to if it is restarting from a crash. In this case, the auto-recovery phase starts which rolls back all pending transactions.
ArcadeDB has different levels of durability based on storage type, configuration and settings.
WAL Flush and Durability
ArcadeDB uses a Write-Ahead Log (WAL) to guarantee transaction durability.
The arcadedb.txWalFlush setting controls whether the WAL is flushed (fsynced) to disk at commit time:
| Value | Behavior | Durability guarantee |
|---|---|---|
|
No flush. WAL is written to the OS page cache but not fsynced. |
Safe against process crashes (OS page cache survives). Not safe against power loss or OS crash - committed transactions may be lost. |
|
Flush without metadata ( |
Safe against power loss. Recommended for production. |
|
Full flush ( |
No additional recovery value over |
Why write order is not persistence order
At commit time ArcadeDB always issues the WAL write() before the data pages are published, so the write-ahead
ordering holds in program order. With txWalFlush=0, however, both writes only land in the OS page cache as dirty
pages: the kernel’s writeback persists dirty pages based on memory pressure, LRU and expiry timers, with no ordering
guarantee between them, and the block layer and the drive cache can reorder further. Ordering on the physical medium
is only ever enforced by an explicit barrier (fsync/fdatasync). On a power loss or kernel panic it is therefore
possible for a data page to have reached the disk while its WAL record did not - despite the WAL having been written
first.
The consequence is worse than losing the most recent transactions: a multi-page transaction can persist some pages and
not others, with no WAL record to replay or repair from - torn records across pages and indexes pointing at data that
never persisted. This is structural corruption, which is why txWalFlush=1 is the production recommendation: it makes
the WAL record durable at commit, after which data pages may lag arbitrarily because recovery replays them from the
WAL.
Two technical notes:
-
txWalFlush=1(fdatasync-class) is sufficient for the WAL’s append-only pattern: POSIXfdatasyncalso flushes the metadata needed to read the data back, which includes the file-size extension of an append. Level2adds only timestamp-class metadata, with no recovery value - benchmarks show no measurable difference between the two. -
txWalFlush=0is safe against process crashes (the OS page cache survives the process), only power loss and kernel panics are at risk.
Data files and WAL deletion
A WAL file is deleted in three places: on a clean close of the database, when the background WAL rotation retires a full WAL file, and at the end of the recovery that replayed it after a crash. Before each deletion ArcadeDB forces to disk the data files that the WAL still protects, because once the WAL is gone the data files are the only copy of those pages.
Only the files that need it are forced: a data file is forced when a page was written to it, or when it was created or
renamed, since its last successful fsync. A file that was only read is skipped, since it holds nothing that is not
already on disk. A file with page writes only is forced with fdatasync, which also persists the file-size change of
an append. A created or renamed file is forced with a full fsync, followed by an fsync of its directory, so its
directory entry survives a power loss too. If an fsync fails, the WAL is kept, and the next pass retries the fsync.
If a clean close cannot complete it, the WAL and the lock file stay on disk, and the next open runs recovery.
| Up to version 26.9.x, a clean close forced every data file of the database, including the files that nothing had written. Each of those fsyncs costs a flush of the device write cache, about half a millisecond per file on a laptop NVMe drive and more on network or cloud block storage. So the close of a database with 100 indexed types took about 100 ms even after a read-only session. Now a close after a read-only session forces no file. After a crash, recovery forces every data file once, because it is not known which writes of the crashed process reached the disk. |
Measured cost
Commit-heavy benchmark (one small insert per transaction, so the WAL flush dominates; Apple silicon laptop, macOS):
| Level | TPS (1 thread) | Latency (1 thread) | TPS (8 threads) | Latency (8 threads) |
|---|---|---|---|---|
|
24,813 |
40 us |
68,772 |
116 us |
|
282 |
3.5 ms |
587 |
13.6 ms |
|
305 |
3.3 ms |
620 |
12.9 ms |
The dramatic gap on macOS is platform-specific: Java’s force() issues F_FULLFSYNC there, which drains the
entire drive write cache (~3 ms on Apple SSDs). On Linux server hardware with NVMe, fdatasync typically costs
30-100 us, so the production penalty is far smaller than this laptop measurement suggests. Reproduce on your own
hardware with mvn -pl engine test -Dtest=WALFlushBenchmark from the ArcadeDB source tree.
|
In production server mode (arcadedb.server.mode=production), ArcadeDB automatically sets txWalFlush=1 if you have not explicitly configured it.
This ensures that production deployments are durable by default.
If you explicitly set txWalFlush=0 in production mode, a warning is logged at startup.
Note the server mode defaults to development, so a server that never sets the mode - and any embedded application -
runs with txWalFlush=0: set the mode (or the setting) explicitly for durability.
|
The WAL flush setting can also be changed per-database or per-thread via the Java API:
// Per-database: every thread's transactions, from now on
database.setWALFlush(WALFile.FlushType.YES_NOMETADATA);
// Per-thread: this thread's transactions, from now on, whatever the database setting says
database.getTransaction().setWALFlush(WALFile.FlushType.YES_FULL);
// Hand the thread back to the database setting
database.getTransaction().setWALFlush(null);
// Per async executor
database.async().setTransactionSync(WALFile.FlushType.YES_NOMETADATA);
setUseWAL() and setAsyncFlush() follow the same rule: on Database they apply to every thread, on
getTransaction() to the calling thread only, and the thread’s own setting wins. Until either is called, the
arcadedb.txWAL and arcadedb.txWalFlush configuration applies.
Before 26.10.1, Database.setWALFlush(), setUseWAL() and setAsyncFlush() changed only the calling thread’s
transactions, despite reading as database-wide. A multi-threaded application that called setWALFlush(YES_FULL) once
at startup had its commits flushed on that thread only (issue #8352). They are database-wide now. An application that
relied on the old per-thread scope, for example a bulk loader turning the WAL off on its own thread while other
threads keep writing, must call the same setter on database.getTransaction() instead.
|
If your storage hardware has battery-backed write cache (BBU/BBWC) or power-loss protection (common in enterprise SSDs and cloud block storage like AWS EBS), txWalFlush=0 is safe even in production because the hardware guarantees that buffered writes reach persistent storage on power loss.
|
Optimistic Transaction
This mode uses the well known Multi Version Control System MVCC by allowing multiple reads and writes on the same records.
The integrity check is made on commit.
If the record has been saved by another transaction in the interim, then an ConcurrentModificationException will be thrown.
The application can choose either to repeat the transaction or abort it.
| ArcadeDB keeps the whole transaction in the host’s RAM, so the transaction size is affected by the available RAM (Heap) memory on JVM. For transactions involving many records, consider to split it in multiple transactions. |
A ConcurrentModificationException always means another transaction wrote the record: it is never raised because of something the transaction did to itself.
In particular, updating a record that the same transaction has already deleted is not a conflict, it simply has no effect - the delete wins, since the record will not exist once the transaction commits.
The order of the two operations does not matter.
Nested transactions and propagation
ArcadeDB does support nested transaction.
If a begin() is called after a transaction is already begun, then the new transaction is the current one until commit or rollback.
When this nested transaction is completed, the previous transaction becomes the current transaction again.