Graph OLAP Engine

The Graph OLAP Engine maintains a read-optimized, columnar representation of your graph alongside the live OLTP data. It uses Compressed Sparse Row (CSR) encoding and flat primitive arrays to deliver 5x–400x speedups on analytical workloads — multi-hop traversals, graph algorithms, and property aggregations — without sacrificing transactional safety.

Why Graph OLAP?

ArcadeDB’s OLTP engine is optimized for point lookups and ACID transactions. Analytical workloads — PageRank, community detection, multi-hop traversals — access millions of edges in tight loops. The row-oriented, pointer-chasing nature of OLTP storage causes cache misses, object overhead, and GC pressure.

The OLAP engine solves this by encoding graph topology as flat int[] arrays and properties as typed columns:

  • Sequential memory access — cache-line friendly, no pointer chasing

  • Zero object allocation — no GC pressure during traversal

  • SIMD-friendly — enables JVM vectorized operations

  • 9x more compact — flat arrays vs. Java object overhead

Graph Analytical View (GAV)

A Graph Analytical View is a named, schema-persisted OLAP snapshot of selected vertex types, edge types, and properties.

GraphAnalyticalView gav = GraphAnalyticalView.builder(database)
    .withName("social")
    .withVertexTypes("Person", "Company")
    .withEdgeTypes("FOLLOWS", "WORKS_AT")
    .withProperties("name", "age", "status")
    .withUpdateMode(UpdateMode.SYNCHRONOUS)
    .build();

Named views are persisted in schema.json and automatically restored on database restart. As of ArcadeDB v26.9.1, a READY view’s CSR snapshot itself can also be persisted to disk at a clean database close and reused on the next open without rescanning the graph; see CSR Persistence.

SQL

Creating a view

CREATE GRAPH ANALYTICAL VIEW social
  VERTEX TYPES (Person, Company)
  EDGE TYPES (FOLLOWS, WORKS_AT)
  PROPERTIES (name, age, status)
  UPDATE MODE SYNCHRONOUS

All clauses after the view name are optional. A minimal view covering the entire graph:

CREATE GRAPH ANALYTICAL VIEW fullGraph

Use IF NOT EXISTS to avoid errors if the view already exists:

CREATE GRAPH ANALYTICAL VIEW IF NOT EXISTS social
  VERTEX TYPES (Person)
  EDGE TYPES (FOLLOWS)
  UPDATE MODE SYNCHRONOUS

You can also materialize edge properties (e.g., weights):

CREATE GRAPH ANALYTICAL VIEW weighted
  VERTEX TYPES (City)
  EDGE TYPES (ROAD)
  EDGE PROPERTIES (distance, toll)
  UPDATE MODE SYNCHRONOUS
  COMPACTION THRESHOLD 50000

Add CCH (…​) to keep a Customizable Contraction Hierarchy per weight property, for fast point-to-point shortest paths. A weight listed there is materialized as an edge property even when EDGE PROPERTIES does not name it:

CREATE GRAPH ANALYTICAL VIEW roads
  EDGE TYPES (ROAD)
  UPDATE MODE SYNCHRONOUS
  CCH (distance, travelTime)

Edge sub-types

Listing an edge type in EDGE TYPES also materializes its sub-types, and a hop on that type answers with the sub-types' edges too, as it does without a view. The view keeps one slice per concrete type, so the order the types are listed in does not matter. Views persisted by an earlier release are rebuilt once.

Altering a view

Change the update mode or compaction threshold of an existing view:

ALTER GRAPH ANALYTICAL VIEW social UPDATE MODE ASYNCHRONOUS
ALTER GRAPH ANALYTICAL VIEW social COMPACTION THRESHOLD 20000

Rebuilding a view

Force a full rebuild of the CSR snapshot:

REBUILD GRAPH ANALYTICAL VIEW social

Dropping a view

DROP GRAPH ANALYTICAL VIEW social
DROP GRAPH ANALYTICAL VIEW IF EXISTS social

Listing views

SELECT FROM schema:graphAnalyticalViews

Builder Options

Method Description Default

withName(String)

Named registration + schema persistence

anonymous

withVertexTypes(String…​)

Filter to specific vertex types

all

withEdgeTypes(String…​)

Filter to specific edge types

all

withProperties(String…​)

Materialize specific vertex properties

all

withEdgeProperties(String…​)

Materialize edge properties (e.g., weights)

none

withUpdateMode(UpdateMode)

OFF, SYNCHRONOUS, or ASYNCHRONOUS

OFF

withCompactionThreshold(int)

Rebuild CSR after N accumulated delta edges

10,000

withContractionHierarchy(String, String…​)

Keep a contraction hierarchy for a weight property, over the given edge types (all of the view’s when none)

none

Async Build for Large Graphs

For large graphs, use buildAsync() to avoid blocking the calling thread:

GraphAnalyticalView gav = GraphAnalyticalView.builder(database)
    .withName("large-graph")
    .withUpdateMode(UpdateMode.ASYNCHRONOUS)
    .buildAsync();

// Wait for build completion
boolean ready = gav.awaitReady(30, TimeUnit.SECONDS);

Update Modes

The GAV supports three synchronization modes between OLTP and OLAP:

Mode Behavior Staleness Use Case

OFF

Marks view STALE on commit; requires manual rebuild

Until rebuild

Batch analytics, static snapshots

SYNCHRONOUS

Applies an overlay on each commit

Zero

Real-time analytics, consistent reads

ASYNCHRONOUS

Triggers background rebuild on commit

Brief BUILDING window

Large graphs, tolerable brief inconsistency

In SYNCHRONOUS mode, the engine captures transaction deltas (new/deleted vertices, added/removed edges, property changes) and merges them into an immutable overlay on top of the base CSR. Readers always see a consistent snapshot via an atomic volatile reference swap.

Writes during a build or rebuild. A view starts tracking changes before it scans the graph, so a transaction that commits while the view is being built or rebuilt is never lost: in SYNCHRONOUS mode it is applied on top of the finished view (checked against what the scan already read, so nothing is counted twice), in OFF mode the view is published STALE, and in ASYNCHRONOUS mode it rebuilds once more. Writers are not blocked by a running build, including REBUILD GRAPH ANALYTICAL VIEW.

In SYNCHRONOUS mode the view being rebuilt keeps serving queries, and keeps absorbing commits, until the rebuild publishes its new CSR: only the caller of REBUILD GRAPH ANALYTICAL VIEW (or of build()) waits for the scan. Since v26.10.1 the scan no longer runs while holding the view’s lock, so neither a committing transaction nor a status check waits behind it (issue #8403); before, the view served during a blocking rebuild lagged the commits until the rebuild finished. A rebuild superseded by a newer one while it scans waits for the newer one and reports its outcome (success or error); a rebuild interrupted by the database closing fails with an error rather than reporting success.

Before 26.10.1, a transaction committed while a view was building could be missing from the view permanently, with the view still reporting itself READY and not stale (issue #8378). Rebuild a view created on an earlier version while writes were running if its answers must match the graph exactly. When the overlay accumulates too many changes (configurable threshold, default 10,000 edges), a background compaction rebuilds the full CSR.

Lightweight edges and bulk loads. A lightweight edge is tracked like any other edge in every mode: its creation and its deletion reach the overlay in SYNCHRONOUS mode, mark an OFF view STALE and rebuild an ASYNCHRONOUS one. Each copy of a duplicated lightweight edge counts as an edge of its own, and a delete removes one copy. Copies have no identity to tell an old one from a new one, so when lightweight edges of a pair change while the view is being built or rebuilt, the view checks that pair against the graph before serving it; if a duplicated copy made them disagree, it rebuilds rather than serve the pair one copy off. Edges loaded with GraphBatch (and so by the graph importer, the HTTP batch endpoint and the gRPC batch load) are written in bulk, without an event per edge, so the view learns of them when the batch closes: OFF goes STALE, ASYNCHRONOUS rebuilds, and SYNCHRONOUS rebuilds too rather than absorbing millions of edges into the overlay. While that rebuild runs the view is not READY, so queries take the ordinary path and never answer without the load. During the batch itself the view still serves what it held before the batch started.

Before 26.11.1 a view never learned about a lightweight edge created after its build, nor about any edge written by GraphBatch: a SYNCHRONOUS view answered as if those edges did not exist, and a transaction or batch that touched nothing else left an OFF view READY (issue #9572). Rebuild a view on such a graph after upgrading.

CSR Persistence

Available since ArcadeDB v26.9.1.

Before v26.9.1, a named view’s CSR was never itself persisted: only its definition (types, properties, update mode) was, in schema.json. Every database open therefore rebuilt the CSR from a full graph scan, and because shutdown() waits for any in-flight rebuild, the cost of that scan landed on close() rather than open() - at 1M vertices, closing a database with a view took over 4 seconds even for a session that ran no query at all.

A READY view (with no pending overlay changes) now writes its CSR to disk at a clean database close, alongside a freshness certificate: the database’s last committed transaction id at the time the CSR was built. On the next open, if a persisted file plausibly applies (see arcadedb.gavPersistCsr below), the view is marked READY immediately, but the file itself is not read yet - see Lazy Restore. That deferred read either:

  • finds the database’s current last committed transaction id still matches the certificate - nothing was committed to the database in between, so the persisted CSR is loaded from disk as-is instead of being rebuilt; or

  • finds a mismatch - anything was committed in between, to any type, not only the ones the view covers - and falls back to the previous behavior: an async rebuild.

This means the speedup only applies to reopening a database that was not written to since its last clean close. A database that receives writes between every session still pays the rebuild on every open, exactly as before v26.9.1.

Persistence is controlled by the arcadedb.gavPersistCsr database-level setting (default true). Set it to false to skip writing the CSR file, for example to avoid its disk footprint or the extra write at close.

Serving queries against a stale CSR while a delta is applied in the background (the way SYNCHRONOUS mode’s overlay already does for live updates) is not part of this: a CSR’s contiguous array layout makes partial coverage expensive to retrofit, so an invalidated certificate always means a full rebuild, not a partial one.

Lazy Restore

Available since ArcadeDB v26.9.1.

A persisted CSR that plausibly applies is not read from disk during open(). Reading and deserializing it is a real, if bounded, cost - up to roughly a second at 10M vertices - and a session that opens a database and never queries the view got nothing for paying it. Instead, open() marks the view READY immediately without touching the file, and the actual read is deferred to whichever of these happens first:

  • a query that actually needs the view’s data; or

  • an explicit call to the view’s awaitReady() method.

Either trigger re-verifies the freshness certificate from scratch at that point (not the one implied at open() time), so a commit that lands after open() but before the first query is still caught correctly and falls back to a rebuild rather than serving stale data. A session that opens and closes a database without ever touching the view now costs what the no-view baseline costs, not what a restore costs - close() does not wait for a restore that was never triggered either.

Since v26.11.1 a query that looks a view up while its restore is in flight (for example the first count push-down, MATCH or shortestPath() after a reopen) waits for the restore instead of taking the slower record path. The wait is bounded by the arcadedb.gavQueryRestoreAwaitTimeout database-level setting (default 5000 ms, 0 disables the wait).

A query that runs while a view is not ready yet (its restore past that wait, its first build still running, or the view stale and not allowed to serve stale results) takes the record path. Since v26.11.1 that is no longer remembered: the OpenCypher plan cached for the query is planned again as soon as the view is ready, so the following runs of the same query use the view.

The arcadedb.gavRestoreAwaitTimeout database-level setting (default 0) forces the wait at open() time instead, for either the deferred restore or a full rebuild: set it to a positive number of milliseconds to trade a slower open() for the view being immediately usable by the query that triggered the reopen, rather than the first query racing (and possibly missing) it.

How CSR Works

The graph topology is stored as two pairs of arrays (forward for outgoing edges, backward for incoming):

Forward CSR (outgoing edges):
  offsets:   [0, 3, 5, 8, ...]     -- one entry per vertex + sentinel
  neighbors: [1, 5, 7, 2, 6, ...]  -- dense neighbor IDs, contiguous per source

  Outgoing neighbors of vertex v = neighbors[offsets[v] .. offsets[v+1])
  Out-degree of vertex v         = offsets[v+1] - offsets[v]   -- O(1)

This layout enables sequential memory access (cache-line friendly) and O(1) degree lookups.

Columnar Property Storage

Properties are stored as typed flat arrays — int[], long[], double[], or dictionary-encoded int[] for strings. Each column has a compact null bitmap (1 bit per vertex). Dictionary encoding maps unique string values to integer codes, achieving near-100% compression for low-cardinality fields.

Edge Properties and Pending Changes

Behavior changed in ArcadeDB v26.9.1.

Edge property columns (EDGE PROPERTIES (…​), used as weights by the shortest-path and spanning-tree algorithms) are laid out in the order the graph had when the view was built. In SYNCHRONOUS mode a commit is absorbed into a pending-changes overlay rather than rebuilding, so from that moment the neighbour list and the columns describe two different orderings.

Before v26.9.1 the view resolved this by reporting no edge properties at all as soon as an overlay existed, and every weighted algorithm silently went back to reading edge records - after every single commit. The view now reconciles the two instead: an edge already in the base graph keeps answering from its column, and an edge the overlay added answers from the value recorded when it was committed. Weighted algorithms therefore keep the fast columnar path while a view is being updated.

Two cases remain that the view cannot answer, and for those the algorithm reads the edge records directly - exact, only slower:

  • several parallel edges join the same two vertices and only some of them were deleted, since the overlay counts deletions per vertex pair and cannot say which of the parallel edges went;

  • a property of an existing edge was changed while another edge of the same type joins the same two vertices, until the rebuild that refreshes the columns completes, since the view cannot tell which of the two changed. Since v26.11.1, a change to an edge that is the only one of its type between its two vertices is kept in the overlay and answered from there, with no rebuild (issue #9437). Adding and deleting edges never have this effect.

Memory Usage

The OLAP representation is significantly more compact than the OLTP equivalent:

  • CSR topology: ~8 bytes per edge (bidirectional)

  • Node ID mapping: ~8 bytes per vertex

  • Columnar properties: 4–8 bytes per vertex per column

  • Null bitmaps: 1 bit per vertex per column

Example: for a graph with 500K vertices and 8M edges, the GAV uses 134.6 MB compared to an estimated ~1.2 GB for the OLTP representation — 9.3x more compact.

long bytes = gav.getMemoryUsageBytes();

Graph Algorithms

The module includes parallelized graph algorithms that operate directly on CSR arrays with zero GC pressure:

Algorithm Description

PageRank

Pull-based, parallel, configurable damping factor and iterations

Connected Components

Parallel union-find for weakly connected components

BFS

Breadth-first search with distance arrays

SSSP (Dijkstra)

Single-source shortest path for weighted graphs

Label Propagation

Community detection

Triangle Counting

Count 3-cliques in the graph

Local Clustering Coefficient

Per-vertex clustering coefficients

GraphAlgorithms algos = new GraphAlgorithms();

// PageRank (20 iterations, damping 0.85)
double[] ranks = algos.pageRank(gav, 20, 0.85);

// Connected Components
int[] components = algos.connectedComponents(gav);

// BFS from a source vertex
int[] distances = algos.bfs(gav, sourceNodeId);

Algorithms and Pending Changes

Behavior changed in ArcadeDB v26.9.1.

In SYNCHRONOUS mode a committed change is absorbed into a pending-changes overlay rather than rebuilding the view. The overlay keeps the slot of a deleted vertex and numbers an added vertex above the base graph, so from that moment the view’s vertex count and the range of its internal vertex numbers stop being the same figure.

Before v26.9.1 the algo.* procedures took them for one and the same, with two consequences:

  • a vertex added after the view was built could be missing from the answer of every procedure, with no error to say so;

  • once the additions outnumbered the deletions, the CSR-accelerated procedures (algo.wcc, algo.pagerank, algo.labelpropagation, algo.localClusteringCoefficient, algo.bfs) failed with an ArrayIndexOutOfBoundsException - one added vertex was enough.

Both are fixed. The vertices a view holds are renumbered onto a gapless range before any algorithm sees them, so every algo.* procedure answers for exactly the vertices that are live at the time of the call - the added ones included, the deleted ones excluded.

While a view is holding pending changes the whole-graph procedures also stop using the parallel kernels listed above, which read the base CSR arrays directly and know nothing of the overlay: they would otherwise answer for the graph as it stood at the last build. Those calls still run off the view’s adjacency, single-threaded rather than parallel, and the compaction that folds the overlay back into the base graph (by default after 10,000 pending edges) restores the parallel path.

Shortest Paths with Contraction Hierarchies (CCH)

A view can keep a Customizable Contraction Hierarchy for a weight property: a precomputed structure that answers point-to-point shortest paths in a time that depends on the depth of the hierarchy rather than on how far apart the two vertices are. It is meant for road, logistics, utility and other infrastructure networks that serve many routing queries.

CREATE GRAPH ANALYTICAL VIEW roads EDGE TYPES (ROAD) UPDATE MODE SYNCHRONOUS CCH (distance)
GraphAnalyticalView roads = GraphAnalyticalView.builder(database)
    .withName("roads")
    .withEdgeTypes("ROAD")
    .withUpdateMode(UpdateMode.SYNCHRONOUS)
    .withContractionHierarchy("distance")
    .build();

Queries go through algo.cch.shortestPath (Cypher) or cchShortestPath() (SQL). A query is answered by the hierarchy when its weight property and relationship types are exactly those of a hierarchy on a ready view; otherwise, and whenever the hierarchy cannot answer for the view’s current state, it is answered by bidirectional Dijkstra. Answers are exact either way, and a transaction holding uncommitted changes always sees them.

How it works

  1. Order - a nested-dissection order of the vertices, computed from the topology alone: small vertex separators (minimum vertex cuts) split the graph recursively, and each separator is ranked above the pieces it separates.

  2. Contraction - eliminating the vertices in that order adds shortcut edges, giving the chordal supergraph of the hierarchy.

  3. Customization - the weight of every edge and shortcut, computed bottom-up from the edge weights. This is the only step that depends on the weights.

  4. Query - the two endpoints climb the hierarchy and meet at the top; shortcuts are then unpacked into the edges they stand for.

On grids, the hardest case for this technique, a query on 360,000 vertices takes about 0.5 ms for a route across the whole grid against about 50 ms for Dijkstra, and the gap grows with the size of the graph. Road networks have much smaller separators than grids and benefit more.

Keeping up with changes

The hierarchy follows the view: every change the view absorbs is followed by a background preparation, on the view’s build executor. In SYNCHRONOUS mode most changes are caught up in milliseconds:

  • A weight change, a closed road (an infinite weight) or a deleted edge re-customizes only the part of the hierarchy above the edges that changed. The view keeps its CSR: the new weight of an edge that is the only one of its type between its two vertices is served from the view’s change overlay, without rebuilding anything.

  • A new edge between two vertices the hierarchy already joins is handled the same way.

  • A new edge between two vertices it does not join, or a new vertex with edges, needs a new contraction. The existing vertex order is kept, with new vertices ranked first, so this costs a contraction and a customization rather than a full build.

  • A weight change on an edge that has a parallel twin of the same type between the same two vertices makes the view rebuild its CSR in the background, because the view cannot tell which of the two changed. The hierarchy then catches up from the rebuilt view.

  • While a preparation is running, queries are answered by bidirectional Dijkstra, so an answer never reflects an older state than the view itself.

  • Directed (OUT/IN) queries are prepared from the start. Undirected (BOTH) ones need a second metric, customized the first time a BOTH query asks for it, so that first query and any during its customization are answered by Dijkstra.

Measured on the California road network (DIMACS USA-road-d.CAL: 1.9M vertices, 4.7M arcs) through a database: a query takes 0.04 to 0.1 ms at any distance, against 18 to 230 ms for Dijkstra; updating 1, 10 and 100 weights takes about 0.7, 8 and 65 ms from commit until the hierarchy answers for it; closing a road about 4 ms; adding a road the hierarchy does not join about 0.9 s; the first build about 13 s. Each query needs a few tens of KB of scratch, sized to the depth of the hierarchy rather than to the graph.

The order is persisted next to the view’s CSR at a clean close (see CSR Persistence) and reused when the database is reopened with nothing committed in between, so a reopen only contracts and customizes.

Suitability

The technique relies on small separators. Graphs without them, such as social graphs or graphs with supernodes, would need a supergraph many times larger than themselves. A hierarchy that would exceed arcadedb.gavCchMaxArcsPerEdge supergraph arcs per edge (default 16) is not built: it reports UNSUITABLE, and queries keep using bidirectional Dijkstra. A REBUILD GRAPH ANALYTICAL VIEW tries again.

Monitoring

SELECT FROM schema:graphAnalyticalViews and GraphAnalyticalView.getStats() list each hierarchy with its status (PREPARING, READY, UNSUITABLE, UNAVAILABLE), size in arcs and memory, how many topologies were built or restored, how many customizations ran, and how many queries it answered or handed back to Dijkstra.

The weight of an edge is the numeric value of its weight property; an edge without one weighs 1, and an edge whose value is negative is not walked.

Query Planner Integration

The Cypher query planner automatically detects ready GAVs and substitutes OLTP traversal operators with CSR-based operators when:

  • A named GAV is registered and in READY state

  • The GAV covers the required vertex and edge types

  • The query does not return edge variables as first-class records (edges in CSR have no RID; edge properties are fully supported)

No query changes are needed — the optimizer transparently accelerates matching traversal patterns.

Relationship uniqueness. Cypher binds a relationship at most once per MATCH clause, so MATCH (a)-[:KNOWS]-(b)-[:KNOWS]-(c) never walks back to a over the edge it just took. A view holds adjacency, not edge records, so since v26.10.1 a hop that could reuse a relationship of another hop in its clause - they share an edge type, or either is untyped - is walked one edge type and direction at a time and remembers which relationship it took, and EXPLAIN marks it unique relationships. Such a clause is served from the view when every one of those hops is anonymous and has a fixed length; a named relationship variable or a variable-length relationship among them sends all of them to the records instead. Hops that cannot collide - (a)-[:KNOWS]→(b)-[:WORKS_AT]→(c) - are unaffected. Before v26.10.1 a view answered such patterns without the rule, counting for example every friend-of-a-friend path that returns to its start over the same edge, and every path that took a self-loop twice (issue #8394). An undirected chain of several hops also counted each self-loop once per adjacency list it sits in, even across separate MATCH clauses.

Views over some of the vertex types. A view built with VERTEX TYPES (…​) holds only the vertices of those types, so the edges that leave them are missing from its adjacency. Edge types do not declare which vertex types they connect, so the planner cannot know that a traversal stays inside the view: since v26.11.1 a view is used for SQL out(), in(), both(), MATCH, TRAVERSE, shortest path, the count push-downs and the algo.* procedures only when it covers every vertex type of the database that holds a vertex, and these statements otherwise read the records. Before that they answered with the covered part of the graph, for example 0 rows for MATCH (b:B)-[:E]→(c:C) on a view over (A, B) (issue #9301). A type with no vertex of its own does not count, because no traversal can reach it: a parent type whose records all live in the sub-types the view lists (Message over Post and Comment), or a type not used yet. The first vertex created in such a type makes the view partial again, without any schema change. A view over a subset of the types still serves a one-hop Cypher scan whose labels it holds. To accelerate the other statements, list every vertex type that holds vertices.

Stale views. A view that is STALE (in OFF mode after a commit, or after a failed rebuild) is not used by the planner unless arcadedb.gavUseWhenStale is true, in which case queries read the snapshot as of its last build. The setting is per database, and applies to every view of that database immediately, including views already built and views restored at open (since v26.10.1; before, only the JVM-wide -Darcadedb.gavUseWhenStale had any effect, issue #7875):

ALTER DATABASE `arcadedb.gavUseWhenStale` true

The builder’s withUseWhenStale(boolean) overrides it for one view, and that override is saved with the view’s definition.

A one-hop pattern feeding an aggregation, such as

MATCH (p:Person)-[:KNOWS]->(f:Person)
WHERE p.city = f.city
RETURN p.city AS city, avg(f.age) AS avgAge, count(*) AS n

is read entirely from the view since v26.10.1: the sources are the view’s nodes carrying the source label, both endpoints read their properties from the view’s columns, and every edge reaches the aggregation as a lightweight row rather than a copied record. EXPLAIN shows it as GAV ONE-HOP SCAN. This applies when the statement is a single MATCH of one fixed-length relationship feeding an aggregating RETURN, with no relationship or path variable and no inline property map, the optimizer would otherwise scan the source label (not seek it through an index), and the view is not stale and covers every vertex type the two labels can match. A count per endpoint alone (RETURN p.city, count(*)) keeps its own, even cheaper, edge-count plan, and an order-sensitive aggregate such as collect(), or a grouped LIMIT/SKIP without ORDER BY, keeps the regular plan.

Counts that build no row. A statement whose only output is count(*) - and a COUNT { } subquery - is answered by a count push-down when its pattern is a chain, a star around one node, or one of the LSQB shapes; EXPLAIN shows it under Using Count Push-Down. Since v26.11.1 the push-downs also take:

  • a property predicate on a node of a chain, a star, the anti-join of LSQB Q8/Q9 or the pair join of LSQB Q2, written inline ((:Message {kind: 'Post'})) or as a WHERE conjunct that reads that node alone (WHERE m.kind = 'Post', WHERE m.created > $since). It is checked once per distinct vertex the position reaches, on the vertex itself, so it costs one record read per such vertex instead of the rows of the whole pattern (issues #9595, #9608). A conjunct that compares two nodes (WHERE a.x < b.x), a property value read off another node ({kind: a.kind}), rand() and pattern predicates still leave the count to the regular plan. A predicate written at one position of a variable holds at every position the variable is written at, and two nodes whose inline maps give one property two different string, boolean or integer values ((c:Message {kind: 'Comment'}), (p:Message {kind: 'Post'})) are known to be different vertices, which proves their relationships different edges the way two unrelated labels do.

  • a MATCH made of parts that share no variable, such as MATCH (a:Person), (b:Person), (c:Person): each part is counted on its own, by its own push-down or plan, and the counts are multiplied. An OPTIONAL MATCH part counts at least one row. Two parts of one MATCH clause are multiplied only when the schema proves their relationships cannot be the same edge, since a clause binds every relationship once. A product that does not fit a 64-bit integer is an error (issue #9596).

  • a WITH that only passes variables on, * and aliases included (WITH message, creator, WITH m AS x), with no aggregation, DISTINCT, WHERE, ORDER BY, SKIP or LIMIT: it leaves the number of rows alone, so the statement is counted as if the clauses on both sides were one query part (issue #9597).

  • count(x) of a node or relationship variable a non-optional MATCH binds, which counts the same rows as count(); a single relationship between two unconstrained nodes, MATCH ()-[e:KNOWS]→() RETURN count(e) (untyped, with several types or undirected too), read from the view’s per-type totals, or without a view in one walk of the edge lists that loads no edge; and MATCH (n) RETURN count(n), the sum of the vertex types' counters (issue #9600). Unlike the SQL SELECT count() FROM KNOWS, which counts edge records, the relationship count includes lightweight edges, which keep no record.

  • the same pattern however it is written: the comma-separated parts of a MATCH are read as the graph they describe, so a chain cut into parts is counted as the chain, and a cycle such as LSQB Q2 (a KNOWS hop closed by a three-hop chain, or the same cycle written as four one-hop patterns starting from the post) is split into the probe hop and the chain by the statistics - the fan-out of each hop sampled on the label it leaves - rather than by the order of the text (issue #9599).

  • an OPTIONAL MATCH only tested for absence (OPTIONAL MATCH (c)-[h:HAS_TAG]→(t) WITH …​ WHERE h IS NULL), which is planned as WHERE NOT (c)-[:HAS_TAG]→(t) and so reaches the anti-join push-down (issue #9598, see OPTIONAL MATCH).

A transaction that holds uncommitted changes - any change, even to types the view does not cover - does not use the view: the view serves the committed graph only, and a query inside the transaction must see the transaction’s own writes. Such a query, Cypher, Gremlin or a graph algorithm, runs on the records until the transaction commits, which PROFILE shows as a plan without the view. Run analytical queries outside write transactions, or after committing, to keep the acceleration.

Lifecycle

// Check status
if (gav.isReady()) { /* safe to query */ }

// Status values: NOT_BUILT, BUILDING, READY, STALE
Status status = gav.getStatus();

// Drop (removes from registry + schema)
gav.drop();

// Shutdown (release resources, schema definition persists)
gav.shutdown();

Benchmark Results

On a graph with 500K vertices and ~8M edges:

Benchmark OLTP OLAP Speedup

1-hop count

6.9 µs

1.2 µs

5.7x

2-hop

101.4 µs

5.1 µs

19.8x

3-hop

1,037 µs

56.4 µs

18.4x

5-hop

194,046 µs

5,141 µs

37.7x

Shortest Path

394 ms/pair

7.5 ms/pair

52.8x

PageRank (20 iter)

124,563 ms

316 ms

394.2x

Connected Components

5,591 ms

197 ms

28.4x

Label Propagation

62,619 ms

645 ms

97.1x

Limitations

  • CSR uses int[] arrays — maximum ~2.1 billion vertices per bucket and ~2.1 billion edges per direction

  • Edges in CSR do not carry their own RID; the Cypher query planner falls back to OLTP only when the query returns an edge variable as a first-class record (e.g., RETURN r). Edge properties are fully supported via withEdgeProperties()

  • Dictionary encoding applies only to string properties

  • Initial build requires a full scan of selected vertex/edge types. A rebuild triggered by an invalidated persisted CSR certificate (or by OFF mode, or by exceeding the SYNCHRONOUS/ASYNCHRONOUS compaction threshold) also requires a full scan - only a clean-close-to-unchanged-reopen cycle skips it