Time Series
ArcadeDB includes a native Time Series engine designed for high-throughput ingestion and fast analytical queries over timestamped data. Unlike bolt-on solutions, the Time Series model is integrated directly into the multi-model core — the same database that stores graphs, documents, and key/value pairs can store and query billions of time-stamped samples with specialized columnar compression, SIMD-vectorized aggregation, and automatic lifecycle management.
Key capabilities:
-
Columnar storage with Gorilla (float), Delta-of-Delta (timestamp), Simple-8b (integer), and Dictionary (tag) compression — 0.4 to 1.4 bytes per sample
-
Shard-per-core parallelism with lock-free writes
-
Block-level aggregation statistics for zero-decompression fast-path queries
-
InfluxDB Line Protocol ingestion for compatibility with Telegraf, Grafana Agent, and hundreds of collection agents
-
Prometheus remote_write / remote_read protocol for drop-in Prometheus backend usage
-
PromQL query language — native parser and evaluator with HTTP-compatible API endpoints
-
SQL analytical functions —
ts.timeBucket,ts.rate,ts.percentile,ts.interpolate, window functions, and more -
Continuous aggregates with watermark-based incremental refresh
-
Retention policies and downsampling tiers for automatic data lifecycle
-
Grafana integration via DataFrame-compatible endpoints (works with the Infinity datasource plugin)
-
Studio TimeSeries Explorer with query, schema inspection, ingestion docs, and PromQL tabs
Creating a TimeSeries Type
Use CREATE TIMESERIES TYPE to define a new time series type.
Every type requires a TIMESTAMP column, zero or more TAGS (low-cardinality indexed dimensions), and one or more FIELDS (high-cardinality measurement values).
CREATE TIMESERIES TYPE SensorReading
TIMESTAMP ts PRECISION NANOSECOND
TAGS (sensor_id STRING, location STRING)
FIELDS (
temperature DOUBLE,
humidity DOUBLE,
pressure DOUBLE
)
SHARDS 8
RETENTION 90 DAYS
COMPACTION_INTERVAL 1 HOURS
Minimal syntax (defaults for everything optional):
CREATE TIMESERIES TYPE SensorReading
TIMESTAMP ts
TAGS (sensor_id STRING)
FIELDS (temperature DOUBLE)
| Option | Default | Description |
|---|---|---|
|
(required) |
Name of the timestamp column |
|
|
Timestamp resolution: |
|
(none) |
Comma-separated |
|
(required) |
Comma-separated |
|
|
Number of shards for parallel writes |
|
(none) |
Automatic deletion of data older than the specified duration (e.g., |
|
(none) |
Cuts sealed blocks at boundaries of this interval. A block that lies inside one query bucket is answered from its statistics without decompressing it, so set it to the bucket width of your most frequent aggregation (see Choosing the layout). Without it blocks are cut by size only and a block spanning several buckets is decompressed |
|
(none) |
Skip creation if the type already exists, instead of failing. The statement returns one row either way, carrying a |
A type stores its columns in the order they are declared, and that order matters: a sample is a positional row, so the column order is what an export records and what a restore reads back.
The TIMESTAMP, TAGS and FIELDS clauses may therefore be written in any order, and TAGS and FIELDS may each
appear more than once. The declaration below creates a type whose columns are f1, ts, t1, f2, in that order:
CREATE TIMESERIES TYPE Interleaved
FIELDS (f1 DOUBLE)
TIMESTAMP ts
TAGS (t1 STRING)
FIELDS (f2 DOUBLE)
The Java schema builder accepts the same orders and keeps them, so a recipe rendered as SQL - which is what the remote schema builder does - creates the same type as the same recipe run against an embedded database. A JSONL export therefore restores into a remote database exactly as it does into an embedded one, whatever order its columns are in.
|
Before 26.10.1 the statement had exactly one slot for each clause, in the order timestamp, tags, fields. A builder recipe that interleaved them was regrouped when rendered as SQL, so the remote builder could create a type whose columns were in a different order than the embedded one - and a JSONL restore of such a type into a remote database was refused rather than silently loading the values into the wrong columns. Every statement written against the old grammar parses and means exactly what it did. |
Altering a TimeSeries Type
Add downsampling policies to automatically reduce resolution of old data:
ALTER TIMESERIES TYPE SensorReading
ADD DOWNSAMPLING POLICY
AFTER 7 DAYS GRANULARITY 1 MINUTES
AFTER 30 DAYS GRANULARITY 1 HOURS
Remove all downsampling policies:
ALTER TIMESERIES TYPE SensorReading DROP DOWNSAMPLING POLICY
A TIMESERIES type stores only its declared TIMESTAMP, TAGS, and FIELDS columns.
It cannot be given a SUPERTYPE, and it cannot be used as a SUPERTYPE for another type: any inherited property
would be silently dropped by every write, so both directions are rejected. For the same reason, CREATE PROPERTY
on a TIMESERIES type is also rejected unless the name matches an already-declared column.
|
Compacting a TimeSeries Type
New samples land in each shard’s mutable bucket, and the maintenance scheduler seals them into compressed blocks every 60 seconds. To seal them now, for example right after a bulk load or before measuring a query on the settled layout, run:
COMPACT TIMESERIES TYPE SensorReading
It runs the same compaction the scheduler runs, requires the updateSchema permission like COMPACT INDEX, and
under HA runs on the leader. The result carries mutableSamplesBefore and mutableSamples, the samples still in the
mutable buckets when it finished: rows appended while it ran stay there until the next pass.
The scheduler (and COMPACT TIMESERIES TYPE) also merges the small blocks a slow feed leaves behind into full-size
ones, so the number of blocks follows the data volume and not the time. A read that is scanning the old blocks at the
very moment they are merged is refused with a retryable error (HTTP 503) instead of returning a partial answer:
run it again. A merge rewrites the sealed file, so it needs free disk space of about the size of that file.
Reads do not need a compaction to be fast. A query for one tag, such as the newest reading of one host, skips every mutable page that holds no row of that tag without reading its rows, whether or not the tail has been sealed yet.
Dropping a TimeSeries Type
DROP TIMESERIES TYPE SensorReading
DROP TIMESERIES TYPE SensorReading IF EXISTS
Ingesting Data
There are four ways to ingest data into a TimeSeries type, listed from fastest to most convenient.
InfluxDB Line Protocol (Recommended for High Throughput)
The fastest remote ingestion path. ArcadeDB exposes an InfluxDB Line Protocol-compatible HTTP endpoint that skips SQL parsing entirely.
POST /api/v1/ts/{database}/write?precision=ns
Content-Type: text/plain
SensorReading,sensor_id=sensor-A,location=building-1 temperature=22.5,humidity=65.0,pressure=1013.25 1708430400000000000
SensorReading,sensor_id=sensor-B,location=building-2 temperature=19.1,humidity=70.0 1708430400000000000
Line Protocol format: <measurement>[,<tag>=<value>…] <field>=<value>[,…] [<timestamp>]
A malformed line is not ingested; the rest of the payload is. When any line was malformed the endpoint answers 400
with a partial-write body: written and dropped counts (a malformed line counts as one dropped sample) and
malformedLines, the 1-based numbers of the rejected lines in the payload (the first 100). A field with no value
(temp, temp= or a line cut short after a field key, such as weather temp=21.5,hum) makes the whole line
malformed. So does a , separator with nothing after it, in the tags or in the fields (weather temp=21.5, or
weather,city=rome, temp=1).
|
Since v26.10.1. Before, a malformed line was only logged and the request answered |
Precision parameter: ns (nanoseconds, default), us (microseconds), ms (milliseconds), s (seconds).
Example with curl:
curl -X POST "http://localhost:2480/api/v1/ts/mydb/write?precision=ns" \
-u root:password \
-H "Content-Type: text/plain" \
--data-binary 'SensorReading,sensor_id=sensor-A temperature=22.5 1708430400000000000
SensorReading,sensor_id=sensor-A temperature=22.6 1708430401000000000'
Example with Python:
import requests
lines = [
"SensorReading,sensor_id=sensor-A temperature=22.5 1708430400000000000",
"SensorReading,sensor_id=sensor-A temperature=22.6 1708430401000000000",
]
requests.post(
"http://localhost:2480/api/v1/ts/mydb/write?precision=ns",
auth=("root", "password"),
headers={"Content-Type": "text/plain"},
data="\n".join(lines),
)
If the type does not exist and auto-creation is enabled (arcadedb.tsAutoCreateType=true), the schema is inferred from the first line: measurement name becomes the type, tags become TAG columns, fields become FIELD columns with inferred types.
Prometheus Remote Write
ArcadeDB acts as a drop-in Prometheus remote storage backend. Configure Prometheus to write to ArcadeDB:
# prometheus.yml
remote_write:
- url: "http://localhost:2480/ts/mydb/prom/write"
basic_auth:
username: root
password: password
The endpoint accepts the standard Prometheus remote_write Protobuf payload (snappy-compressed).
Each time series is mapped to an ArcadeDB TimeSeries type named after the name label.
Types are auto-created if they do not exist.
(Since v26.11.1) A label with an empty value is the same as an absent label, as in Prometheus: it is stored as absent and {label=""} selects it. A series carrying a label that the existing type does not declare as a tag is dropped and the request answers 400 naming the label (the other series of the request are stored); Prometheus does not retry a 400. Set arcadedb.timeSeriesUndeclaredKeys=ignore to discard the extra label instead.
|
(Since v26.10.1) if your collector or proxy sends an Before, on Genuine retries still work as before: re-sending the identical body under the same id is applied only once. |
SQL INSERT
Standard ArcadeDB SQL syntax works for TimeSeries types:
-- Single row
INSERT INTO SensorReading
SET ts = date('2026-02-20 10:00:00', 'yyyy-MM-dd HH:mm:ss'),
sensor_id = 'sensor-A',
location = 'building-1',
temperature = 22.5,
humidity = 65.0
-- Batch insert
INSERT INTO SensorReading
(ts, sensor_id, location, temperature, humidity)
VALUES
(date('2026-02-20 10:00:00', 'yyyy-MM-dd HH:mm:ss'), 'sensor-A', 'building-1', 22.5, 65.0),
(date('2026-02-20 10:00:01', 'yyyy-MM-dd HH:mm:ss'), 'sensor-A', 'building-1', 22.6, 64.8),
(date('2026-02-20 10:00:02', 'yyyy-MM-dd HH:mm:ss'), 'sensor-B', 'building-2', 19.1, 70.0)
-- CONTENT syntax
INSERT INTO SensorReading
CONTENT { "ts": 1771581600000, "sensor_id": "sensor-A", "temperature": 22.5 }
(Since v26.11.1) INSERT fails if it sets a property that is not a declared tag or field of the type (for example a misspelled tag), because the sample would otherwise be filed under a different series. Set arcadedb.timeSeriesUndeclaredKeys=ignore on the database to discard such properties instead. The same setting governs line protocol, gRPC and Prometheus writes.
|
The timestamp column of a TimeSeries type is stored as a number, so INSERT does not accept an ISO-8601 string for it (Cannot convert type String to LONG): write an epoch value (as in the CONTENT example, milliseconds for a type with PRECISION MILLISECOND) or build it with the date() function, which interprets the text in the server’s time zone. ISO-8601 strings are accepted in WHERE conditions on the timestamp column.
|
Java Embedded API
The fastest path — bypasses all protocol and SQL overhead:
TimeSeriesEngine engine = database.getSchema()
.getTimeSeriesType("SensorReading").getEngine();
long[] timestamps = { 1708430400000000000L, 1708430401000000000L };
String[] sensorIds = { "sensor-A", "sensor-A" };
double[] temperatures = { 22.5, 22.6 };
database.transaction(() -> {
engine.appendSamples(timestamps,
new Object[] { sensorIds, temperatures });
});
Appends and Transactions
A TimeSeries append is not part of the transaction that contains it. This is true of every ingestion method above -
SQL INSERT, the REST write endpoint, the line protocol, Prometheus remote write and the Java embedded API alike:
-
the samples are committed as they are appended, so they are durable and visible to every other connection as soon as the append returns, before you commit anything;
-
rolling the enclosing transaction back does not remove them.
Use a session or a transaction around an append for what it does do - it holds that session’s lock and principal and keeps its idle timer from reaping it while you ingest - and not for atomicity with the samples.
|
The reason is the shard’s own append lock. Every append writes the shard’s header page, which holds the sample count and the min/max timestamps, so two appends to one shard always want the same page version; serializing them behind a lock the append itself releases at its own commit is what keeps them from conflicting. An append that wrote through the caller’s transaction could not be serialized that way, because the commit would belong to the caller and a shard cannot hold that lock until an arbitrary user transaction ends without blocking every other writer of that shard. Both would then stage the same page and one would lose its whole transaction to a concurrent-modification error that is not the shard’s to retry. |
Ingestion Method Comparison
| Method | Throughput | Overhead | Best For |
|---|---|---|---|
Java Embedded API |
~0.5-1 us/sample |
None (direct) |
Embedded applications |
InfluxDB Line Protocol |
~1-5 us/sample |
Text parsing |
Remote ingestion, Telegraf |
Prometheus Remote Write |
~2-10 us/sample |
Protobuf + Snappy |
Prometheus ecosystems |
SQL INSERT |
~50-100 us/sample |
SQL parsing + planning |
Ad-hoc inserts, small batches |
Querying Time Series Data
SQL Queries
Time series types support standard SQL SELECT with WHERE, GROUP BY, and ORDER BY.
Time range conditions (BETWEEN, >, >=, <, <=, =) on the timestamp column are pushed down to the storage engine for efficient range scans. A time range on one branch of an OR never narrows the other branch, and two equalities on the same tag in one AND match nothing when the values differ, exactly as on a document type.
-- Basic range query
SELECT ts, sensor_id, temperature, humidity
FROM SensorReading
WHERE ts BETWEEN '2026-02-19' AND '2026-02-20'
AND sensor_id = 'sensor-A'
ORDER BY ts
-- Aggregation with time bucketing
SELECT ts.timeBucket('1h', ts) AS hour,
sensor_id,
avg(temperature) AS avg_temp,
max(temperature) AS max_temp,
min(temperature) AS min_temp,
count(*) AS sample_count
FROM SensorReading
WHERE ts BETWEEN '2026-02-19' AND '2026-02-20'
GROUP BY hour, sensor_id
ORDER BY hour
An aggregation that groups by the time bucket, by up to four TAG columns, or by both is answered by the storage engine in one pass (AGGREGATE FROM TIMESERIES in EXPLAIN), the shape of a per-host hourly report. Every projected tag must also be in the GROUP BY, and the aggregates are avg, min, max, sum and count(*) over numeric fields. A query outside that shape (a field in the GROUP BY, a projected tag that is not grouped by, count(field), DISTINCT) gives the same answer through the generic row-at-a-time plan, which is slower.
TimeSeries SQL Functions
ArcadeDB provides a comprehensive set of ts.* SQL functions for time series analytics.
| Function | Description |
|---|---|
|
Truncates a timestamp to the nearest interval boundary for |
|
Returns the value corresponding to the earliest timestamp in the group. |
|
Returns the value corresponding to the latest timestamp in the group. |
|
Per-second rate of change. Optional 3rd parameter ( |
|
Difference between the last and first values in the group. |
|
Moving average with a configurable window size. |
|
Gap filling. Methods: |
|
Pearson correlation coefficient between two series. |
|
Approximate percentile calculation (0.0-1.0). E.g., |
|
Window function: returns the value from a previous row. |
|
Window function: returns the value from a subsequent row. |
|
Window function: sequential 1-based row numbering. |
|
Window function: rank with ties, gaps after ties. |
Moving the bucket grid
By default buckets are multiples of the interval counted from the Unix epoch, which was a Thursday at 00:00 UTC: a
'1w' bucket starts on Thursday and a '1d' bucket starts at 08:00 in UTC+8. The optional third parameter of
ts.timeBucket moves the grid; use one of:
-- weeks start on Monday
SELECT ts.timeBucket('1w', ts, {'origin': '2024-01-01T00:00:00Z'}) AS week, sum(kwh) FROM Energy GROUP BY week
-- days start at local midnight in UTC+8
SELECT ts.timeBucket('1d', ts, {'offset': '-8h'}) AS day, sum(kwh) FROM Energy GROUP BY day
SELECT ts.timeBucket('1d', ts, {'timezone': '+08:00'}) AS day, sum(kwh) FROM Energy GROUP BY day
-
origin: any instant a bucket starts at (a bare instant as third parameter means the same). -
offset: a signed duration added to the epoch grid. -
timezone: local midnight (local Monday for'1w') of a zone with a fixed offset such as'+08:00'. Zones with daylight saving are refused, because their days do not all have the same length.
The same grid is used by the aggregation push-down, the HTTP/gRPC/Grafana bucketOrigin, continuous aggregates
(write the option in the defining query) and downsampling tiers, which take an OFFSET clause:
ALTER TIMESERIES TYPE SensorReading ADD DOWNSAMPLING POLICY AFTER 30 DAYS GRANULARITY 1 DAYS OFFSET -8 HOURS
Blocks already downsampled are not re-bucketed if you change the offset later.
Examples:
-- Rate of change with counter reset detection
SELECT ts.timeBucket('5m', ts) AS window,
ts.rate(request_count, ts, true) AS requests_per_sec
FROM HttpMetrics
WHERE ts > '2026-02-20T10:00:00Z'
GROUP BY window
-- Percentile calculation
SELECT ts.timeBucket('1h', ts) AS hour,
ts.percentile(latency_ms, 0.99) AS p99,
ts.percentile(latency_ms, 0.50) AS median
FROM ServiceMetrics
GROUP BY hour
-- Gap filling with linear interpolation
SELECT ts.timeBucket('1m', ts) AS minute,
ts.interpolate(temperature, 'linear', ts) AS temp
FROM SensorReading
WHERE ts BETWEEN '2026-02-20T10:00:00Z' AND '2026-02-20T11:00:00Z'
GROUP BY minute
-- Correlation between two fields
SELECT ts.correlate(temperature, humidity) AS correlation
FROM SensorReading
WHERE ts BETWEEN '2026-02-19' AND '2026-02-20'
Dedicated JSON Query Endpoint
A simplified REST endpoint is available for programmatic access and Grafana integration:
POST /api/v1/ts/{database}/query
Content-Type: application/json
{
"type": "SensorReading",
"from": "2026-02-19T00:00:00Z",
"to": "2026-02-20T00:00:00Z",
"columns": ["temperature", "humidity"],
"tags": { "sensor_id": "sensor-A" },
"aggregation": "AVG",
"bucketInterval": "1h"
}
For raw (non-aggregated) queries, omit the aggregation and bucketInterval fields.
The optional bucketOrigin (epoch milliseconds, also available on the Grafana endpoint, on gRPC as bucket_origin_ms
and in the Java client as TimeSeriesQuery.bucketOrigin()) moves the bucket grid like the origin option of
ts.timeBucket, see Moving the bucket grid.
from, to, bucketInterval and maxDataPoints must be whole numbers. A value with a fractional part
is refused with 400 and a message naming the member, rather than being rounded down to something you did not
ask for. A bucket width computed by division therefore has to be rounded by the client, deliberately.
|
To retrieve the most recent data point:
GET /api/v1/ts/{database}/latest?type=SensorReading&tag=sensor_id:sensor-A
Repeat tag once per tag column to narrow to a single series; every occurrence must match:
GET /api/v1/ts/{database}/latest?type=SensorReading&tag=sensor_id:sensor-A&tag=region:eu
|
Since 26.10.1, a tag name that is not a On This applies to |
Missing Values
A sample can be absent, which means "no measurement here" rather than a value of zero. You write one either as
NaN or as null, and both are stored the same way on a DOUBLE or FLOAT column. Aggregations treat it that
way:
-
MIN,MAX,SUMandAVGskipNaNsamples. Over a window that holds at least one real value they answer from the real values alone:SUMis the sum of the real samples andAVGdivides it by the number of real samples, not by the number of samples. Over a window whose samples are allNaNthey returnNaN, meaning the window has no measurement to report. -
count(*)counts rows, so it countsNaNsamples too. In SQL,count(field)counts only the rows where the field is not null, as on any other type, and returns aLong. -
A genuine
Infinityor-Infinitysample is a measurement and is returned as such. -
Downsampling averages the real samples of each bucket; a bucket with no real sample is downsampled to
NaN.
This is the semantics of SQL’s NULL under SUM, AVG and count(*). The PromQL endpoint is the one exception:
its sum, avg, sum_over_time and avg_over_time propagate NaN, because that is what Prometheus does and a
PromQL query is expected to answer what Prometheus would.
Over the REST and Grafana endpoints an absent value is serialized as JSON null, since JSON has no NaN literal.
A Grafana panel therefore draws a gap for such a bucket rather than a dip to zero. Over SQL it reads back as NULL,
both as a projected column and as the result of an aggregate with nothing to aggregate:
SELECT value FROM Sensor -- an absent sample is NULL
SELECT avg(value) FROM Sensor GROUP BY bucket -- a bucket with no real sample is NULL
A NaN that arithmetic itself produces over real samples - the sum of +Infinity and -Infinity, for instance -
is an undefined total rather than an absence, and is returned as NaN.
|
Only a |
|
Before 26.10.1 a |
|
Before 26.10.1 |
|
Since 26.10.1 each sealed block records an identity of its own, which is what lets a long-running read - a wide
query, a PromQL range, an The sealed store format version moved to 2 for this, so a store written by 26.10.1 cannot be opened by an earlier version - deliberately, since the alternative is an older version reading it as if it held nothing. In a cluster, keep every node on the same version. |
|
Before 26.9.1 an all- |
Which Columns Can Be Aggregated
SUM, AVG, MIN and MAX read the column’s values, so they need a column that is stored as a number. That
means a numeric FIELD: DOUBLE, FLOAT, LONG, INTEGER, SHORT, BYTE, BOOLEAN (as 1 and 0), and the
datetime types.
A column that is not stored as a number is refused, and the message names it:
{ "error": "Aggregation SUM cannot be applied to column 'host': a TAG column of type STRING is not stored as a number" }
Tags are always stored as text, whatever type they are declared with, so a tag cannot be aggregated even when it
is declared LONG. The same applies to a STRING field and to the timestamp column itself.
COUNT is the exception and works on any column. The SQL push-down also behaves differently: rather than refusing
the query it simply falls back to the generic aggregation path, so SELECT … max(host) … GROUP BY and
count(field) … GROUP BY (which skips null values) still answer. Only count(*) is answered by the push-down.
|
Before 26.10.1 such a request was not refused, it was answered wrongly: the recent samples were read as a column
of zeros, so The same fix corrected When upgrading: a |
Aggregating from Embedded Java
Aggregation from embedded code goes through TimeSeriesEngine.aggregateMulti, which reads any number of columns
in a single pass:
TimeSeriesEngine engine = database.getSchema()
.getTimeSeriesType("SensorReading").getEngine();
List<ColumnDefinition> columns = ((LocalTimeSeriesType) database.getSchema()
.getType("SensorReading")).getTsColumns();
// The position of 'temperature' inside an engine row, from its position in the schema.
int temperature = TimeSeriesGateway.aggregationRowIndex(columns,
TimeSeriesGateway.findColumnIndex("temperature", columns));
MultiColumnAggregationResult result = engine.aggregateMulti(fromTs, toTs,
List.of(new MultiColumnAggregationRequest(temperature, AggregationType.AVG, "avg_temp")),
3_600_000L, null);
for (long bucket : result.getBucketTimestamps())
System.out.println(bucket + " -> " + result.getValue(bucket, 0));
|
When upgrading to 26.10.1: before this release the two halves of the aggregation engine disagreed about this
number, so a type whose The single-column |
PromQL Query Language
ArcadeDB includes a native PromQL parser and evaluator, providing Prometheus-compatible query capabilities without requiring an external Prometheus server.
PromQL via HTTP API
The PromQL endpoints follow the Prometheus HTTP API format and return standard {status: "success", data: {…}} JSON responses.
Instant query:
GET /ts/{database}/prom/api/v1/query?query=avg(cpu_usage{host="srv1"})&time=1700000000
The time parameter is Unix seconds (float). Defaults to current time if omitted.
Range query:
GET /ts/{database}/prom/api/v1/query_range?query=rate(http_requests_total[5m])&start=1700000000&end=1700003600&step=60
All timestamps in Unix seconds. The step parameter accepts a duration string (60s, 1m) or seconds as a float.
|
Since v26.9.1 Earlier versions accepted such values and could take a very long time to answer, or not answer at all, on an extremely wide range. |
|
Since v26.10.1 instant and range queries read samples in a streaming fashion instead of loading the whole lookback window into memory first. A query over a wide time range now uses memory proportional to the number of series returned, not to the number of samples scanned to compute it. |
Label discovery:
GET /ts/{database}/prom/api/v1/labels
GET /ts/{database}/prom/api/v1/label/{name}/values
GET /ts/{database}/prom/api/v1/label/{name}/values?start=1700000000&end=1700003600
GET /ts/{database}/prom/api/v1/series?match[]=cpu_usage{host=~"srv.*"}
These endpoints follow the Prometheus data model, where a label with an empty value counts as absent. A tag
that holds no value for some samples is therefore not offered as an empty entry by label/{name}/values, and
series does not report a series whose only difference from another is an empty label. The generic
/api/v1/ts/{database}/query endpoint is unaffected and still reports exactly what the samples hold.
|
|
Since v26.10.1 Both are optional. A request sending neither covers the whole series and, for Earlier versions ignored both parameters and always answered over the whole retention period, so a Grafana variable scoped to the last hour was offered every value the label had ever held.
|
Supported PromQL Features
| Category | Supported |
|---|---|
Vector selectors |
|
Range selectors |
|
Aggregations |
|
Rate functions |
|
Over-time functions |
|
Math functions |
|
Label functions |
|
Other |
|
Operators |
|
PromQL via SQL
The promql() SQL function calls the PromQL evaluator from within SQL queries:
-- Instant query at explicit time
RETURN promql('cpu_usage{host="srv1"}', 1700000000000)
-- Rate calculation
RETURN promql('rate(http_requests_total[5m])')
-- Scalar arithmetic
RETURN promql('2 + 3 * 4', 1000)
The function accepts 1-2 arguments: the PromQL expression (required) and an optional evaluation timestamp in milliseconds.
Prometheus Remote Read
Configure Prometheus to read from ArcadeDB for long-term storage:
# prometheus.yml
remote_read:
- url: "http://localhost:2480/ts/mydb/prom/read"
basic_auth:
username: root
password: password
The endpoint accepts the standard Prometheus remote_read Protobuf query (snappy-compressed) and supports =, !=, =~, !~ label matchers.
Continuous Aggregates
Continuous aggregates are pre-computed time-bucketed rollups that are automatically refreshed when new data is inserted. They dramatically speed up common dashboard queries by maintaining materialized summaries.
-- Create a continuous aggregate
CREATE CONTINUOUS AGGREGATE hourly_temps AS
SELECT ts.timeBucket('1h', ts) AS hour,
sensor_id,
avg(temperature) AS avg_temp,
max(temperature) AS max_temp,
count(*) AS cnt
FROM SensorReading
GROUP BY hour, sensor_id
-- Query the aggregate like any other type
SELECT * FROM hourly_temps
WHERE hour BETWEEN '2026-02-19' AND '2026-02-20'
-- Manual refresh
REFRESH CONTINUOUS AGGREGATE hourly_temps
-- Drop
DROP CONTINUOUS AGGREGATE hourly_temps
The defining query must reference a TimeSeries source type and contain a ts.timeBucket() call with a GROUP BY clause.
After creation, every committed insert into the source type triggers an incremental refresh using watermark tracking — only new data since the last watermark is processed.
Inspect continuous aggregates via:
SELECT FROM schema:continuousAggregates
This returns name, query, source type, bucket column, bucket interval, watermark timestamp, watermarkSet, status, and metrics for each aggregate.
watermarkSet tells a watermark of 0 that has actually been reached — an aggregate whose newest bucket is the epoch — apart from an aggregate that has never been refreshed, which reports 0 as well.
The watermark is the bucket boundary of the newest row the last refresh produced.
Each refresh deletes the backing rows at or after it and recomputes them from the source, so the incomplete bucket that was open when the previous refresh ran is always rebuilt rather than duplicated.
A defining query may carry its own WHERE clause, including a disjunction: the watermark filter is added as a separate conjunct around it, so WHERE a OR b keeps meaning ts >= watermark AND (a OR b).
Retention Policies
Retention policies automatically delete data older than a specified duration. Set the retention period during type creation:
CREATE TIMESERIES TYPE SensorReading
TIMESTAMP ts
TAGS (sensor_id STRING)
FIELDS (temperature DOUBLE)
RETENTION 90 DAYS
A background maintenance scheduler (60-second interval) automatically enforces retention and downsampling policies on all TimeSeries types.
Downsampling Policies
Downsampling reduces the resolution of old data to save storage while preserving long-term trends. Multiple tiers can be defined to progressively reduce resolution as data ages:
ALTER TIMESERIES TYPE SensorReading
ADD DOWNSAMPLING POLICY
AFTER 7 DAYS GRANULARITY 1 MINUTES
AFTER 30 DAYS GRANULARITY 1 HOURS
In this example, data older than 7 days is downsampled to 1-minute resolution, and data older than 30 days is further reduced to 1-hour resolution. The downsampling process aggregates values using AVG within each granularity bucket, preserving tag groupings.
Reads That Cross a Downsampling Pass
Downsampling replaces samples rather than deleting them, so a read already in flight when a pass runs cannot finish coherently: the rows it has already returned are the original samples, and the rows still to come are their coarser replacements. Joining the two would report the bucket at the crossing point at two different resolutions.
Such a read is therefore refused rather than allowed to return a short answer:
-
Over HTTP — including PromQL and the Grafana endpoints — the response is
503, so a client that already retries on503recovers on its own. -
Over gRPC, Bolt, MongoDB and Postgres the equivalent retryable status is returned.
-
EXPORT DATABASEfails only the TIMESERIES type it was reading; every other type in the database is still exported.
Re-running the read succeeds and returns a complete answer at a single resolution, because downsampling is a scheduled maintenance event rather than something that happens per request.
A retention pass is different. It deletes old data outright, so a read that crosses one still completes and simply returns the rows that remain.
What an Export Reports
The EXPORT DATABASE summary carries two counters for TIMESERIES data that changed underneath it:
| Field | Meaning |
|---|---|
|
Blocks a retention pass deleted while the export was reading them. Their samples are genuinely gone, so the export is still valid for the database as it now stands — but it is not the snapshot it would have been a moment earlier. Reported, never fatal. |
|
Types whose export was cut short by a downsampling pass. The rows already written for such a type are real but incomplete, and the export as a whole is reported as failed so the gap cannot pass unnoticed. Re-run the export for a complete result. |
Grafana Integration
ArcadeDB provides Grafana DataFrame-compatible HTTP endpoints that work with the Grafana Infinity datasource plugin — no custom plugin is needed.
Endpoints:
| Endpoint | Description |
|---|---|
|
Datasource health check |
|
Discovers TimeSeries types, fields, tags, and available aggregation types |
|
Multi-target query returning Grafana DataFrame wire format |
The query endpoint supports raw queries, aggregated queries (SUM/AVG/MIN/MAX/COUNT), tag filtering, field projection, and automatic bucket interval calculation from Grafana’s maxDataPoints setting.
Grafana Infinity datasource configuration:
-
Install the Grafana Infinity datasource plugin.
-
Add a new Infinity datasource.
-
Set the base URL to
http://<arcadedb-host>:2480. -
Configure authentication (Basic Auth with ArcadeDB credentials).
-
Use the health endpoint for health checks.
Studio TimeSeries Explorer
The ArcadeDB Studio web interface includes a dedicated TimeSeries Explorer accessible from the main navigation sidebar.
The explorer provides four tabs:
-
Query — Time range selector, aggregation controls, field checkboxes, interactive charts (ApexCharts with zoom), data table with pagination, auto-refresh capability
-
Schema — Type introspection with column roles (TIMESTAMP/TAG/FIELD badges), diagnostics (total samples, shards, time range), configuration details, downsampling tiers, per-shard statistics
-
Ingestion — Documentation and examples for all four ingestion methods with a method comparison table
-
PromQL — PromQL expression input, instant/range toggle, time controls, chart and table rendering of results
HTTP API Reference
| Method | Endpoint | Description |
|---|---|---|
|
|
InfluxDB Line Protocol ingestion. Body: plain text lines. Returns 204 on success. |
|
|
JSON query endpoint. Supports raw and aggregated queries with tag filtering and field projection. A |
|
|
Returns the most recent data point, with optional tag filter. Repeat |
|
|
PromQL instant query. |
|
|
PromQL range query. |
|
|
List all label names (metric names and tag columns). |
|
|
List distinct values for a label. |
|
|
Find series matching PromQL selector(s). |
|
|
Prometheus remote_write endpoint (snappy-compressed Protobuf). |
|
|
Prometheus remote_read endpoint (snappy-compressed Protobuf). |
|
|
Grafana datasource health check. |
|
|
Grafana metadata discovery. |
|
|
Grafana DataFrame query endpoint. |
High Availability
In an HA cluster, TimeSeries data is handled as follows:
-
Mutable bucket data (
.tstbfiles) is replicated to all followers via ArcadeDB’s standard page-replication protocol -
Sealed store data (
.ts.sealedfiles) is replicated too, but by its own path: compaction, retention and downsampling run on the leader only, and the rewritten sealed file is shipped to the followers together with the change that clears the matching mutable data. Followers do not compact independently -
Reads are consistent after failover: the engine queries both sealed and mutable layers, so a newly promoted leader returns correct results even before its first compaction cycle
-
Compaction lag: a node whose maintenance scheduler has not run yet (default: 60 seconds) serves the same results, just with more of the answer coming from the mutable layer
-
Backups and snapshots taken on a follower are consistent. Because the sealed file and the clearing of the mutable data arrive together, a copy taken across that moment used to be able to capture both and restore with duplicated samples; since 26.10.1 the copy waits for the install to finish
Integrity Checking
|
Starting in 26.9.1, if a type’s sealed store fails to open at all (for example, a corrupted |
A TimeSeries type has neither record buckets nor indexes, so CHECK DATABASE walks it with a
pass of its own that reads all three of the on-disk formats it owns:
-
the mutable bucket (
.tstb) — page 0’s counters are reconciled against the samples its data pages actually hold -
the tag dictionary (
.tstd) — the declared entry count against the entries walked -
the sealed store (
.ts.sealed) — the header, the block directory, the block bounds, and the CRC32 of every block
CHECK DATABASE
CHECK DATABASE TYPE SensorReading
The CRC32 pass reads the whole sealed file. That is deliberate: a sealed block is otherwise verified only on the first read that touches it, so a block nothing queries is a block nothing verifies.
The DEEP tier
CHECK DATABASE DEEP adds the checks that have to decompress the blocks. A CRC proves the bytes are the bytes that
were written; it proves nothing about whether they mean what the block claims. Three things the read paths answer
queries from, without ever looking at a value, are verified only here:
| Claim | What a wrong one does |
|---|---|
Timestamps are sorted |
The range iterator binary-searches them, so a query over the block silently returns a subset of the matching rows |
Per-column |
Aggregation push-down answers |
Distinct tag values |
Block pruning skips a whole block whose declaration omits the value being filtered on, hiding those rows entirely |
Each of those produces a wrong answer rather than an error, which is why nothing else in the engine reports them.
|
|
CHECK DATABASE DEEP
CHECK DATABASE TYPE SensorReading FIX DEEP
What FIX repairs
FIX repairs only what is derived from the data, and never a sample:
-
the mutable bucket’s page-0 counters (sample count, min and max timestamp)
-
the sealed header’s block count and global timestamp bounds
-
the tail of a sealed append that did not complete
None of those is cosmetic. The global bounds are read straight out of the header when the file is opened rather than recomputed, so a range query pruned against a wrong bound silently misses data the file holds; and because new blocks are appended at the end of the file, a tail no reader can see also hides every block appended after it.
A sealed block that fails its CRC, or that DEEP finds inconsistent, is reported and left untouched. It is the only
copy of the samples in it, so discarding it is an operator’s decision - and in an HA cluster the sealed store is
derived, so a node can rebuild one by recompacting from its replicated mutable pages.
The report carries totalTimeSeriesTypes, totalTimeSeriesShards, totalTimeSeriesSamples and
totalTimeSeriesSealedBlocks alongside the findings, so a clean result can be told apart from a pass that never ran.
Findings appear in warnings with the offending type named in corruptedTimeSeries; repairs are counted in
repairedTimeSeries and listed in timeSeriesRepairs.
Comparison with Other TimeSeries Databases
| Feature | ArcadeDB | InfluxDB 3 | TimescaleDB | Prometheus | QuestDB |
|---|---|---|---|---|---|
License |
Apache 2.0 |
MIT (core) |
Apache 2.0 (core) |
Apache 2.0 |
Apache 2.0 |
Multi-Model |
Graph + Document + K/V + TimeSeries + Vector in one engine |
TimeSeries only |
Relational + TimeSeries (PostgreSQL extension) |
Metrics only |
TimeSeries only (SQL) |
Query Languages |
SQL, PromQL, Cypher, Gremlin, GraphQL, MongoDB QL |
SQL, InfluxQL |
Full PostgreSQL SQL |
PromQL |
SQL (PG wire) |
Ingestion Protocols |
Line Protocol, Prometheus remote_write, SQL, Java API |
Line Protocol, SQL |
SQL (INSERT, COPY) |
Prometheus scrape, remote_write |
Line Protocol, SQL, CSV |
Compression |
Gorilla, Delta-of-Delta, Simple-8b, Dictionary |
Parquet native (Delta, Dict, Snappy/ZSTD) |
Gorilla, Delta-of-delta, Simple-8b, Dictionary, LZ4 |
Gorilla (~1.37 B/sample) |
ZFS-level + Parquet for cold tier |
Continuous Aggregates |
Yes (watermark-based, auto-refresh on commit) |
Materialized views |
Yes (policy-based refresh) |
Recording rules |
No |
Downsampling |
Yes (multi-tier, automatic) |
Via compaction |
Via continuous aggregates + retention |
Recording rules |
No |
Retention Policies |
Yes (automatic, per-type) |
Yes |
Yes (per-chunk) |
Yes (per-block) |
Yes (per-partition) |
Grafana Integration |
DataFrame endpoints (Infinity plugin) |
Native plugin |
PostgreSQL datasource |
Native plugin |
PostgreSQL datasource |
PromQL Support |
Native (parser + evaluator + HTTP endpoints) |
No |
Via adapter |
Native |
No |
Embeddable (in-process) |
Yes (Java library) |
No |
No (requires PostgreSQL) |
No |
No (separate process) |
SIMD Aggregation |
Yes (Java Vector API) |
Via DataFusion (Arrow) |
No |
No |
Yes (AVX2) |
Prometheus Remote Write/Read |
Yes |
No |
Via adapter |
Native |
No |
Graph + TimeSeries Queries |
Yes (native cross-model) |
No |
No |
No |
No |
Choosing the layout: COMPACTION_INTERVAL and SHARDS
For a type that is bulk-loaded and then aggregated in fixed buckets (for example hourly averages), declare
COMPACTION_INTERVAL with the bucket width of that aggregation and leave SHARDS at its default:
CREATE TIMESERIES TYPE Point TIMESTAMP ts TAGS (host STRING)
FIELDS (uu DOUBLE, us DOUBLE, ui DOUBLE)
COMPACTION_INTERVAL 1 HOURS
-
A sealed block answers an aggregate from its min/max/sum statistics only when it lies inside ONE query bucket. Without an interval a block covers many hours and is decompressed (
PROFILEreports it underslow); with1 HOURSthe blocks are cut at the hour and are counted underfast. -
Other queries on the same type are not penalised. A last-point query (
ORDER BY ts DESC LIMIT 1) walks the blocks newest-first and stops at the first block that has the row, and a finer bucket (timeBucket('1m', ts)) over a short range reads only the blocks of that range, so smaller blocks only mean less to decompress. Ingest is not slowed either: blocks are cut when the mutable data is compacted, not while samples are appended. -
The price is more, smaller blocks (each shard cuts at every interval boundary), so a very small interval on a long retention multiplies the block count. Pick the bucket of the aggregation you run most, not the smallest one you might ever run.
-
SHARDStrades ingest parallelism for fragmentation: every shard cuts its own blocks, so doubling the shards roughly doubles the blocks per interval, and queries merge one more stream. The default (available cores minus 1) suits concurrent writers; for a type loaded by one writer and queried more than written, fewer shards are enough. -
Run
COMPACT TIMESERIES TYPE <name>after a bulk load: mutable samples are not in sealed blocks yet, and only sealed blocks are answered from their statistics.
Architecture
The TimeSeries engine uses a two-layer storage architecture:
-
Mutable bucket — An append-only in-memory buffer backed by ArcadeDB’s
PaginatedComponent. New samples land here first. This layer is ACID-transactional and replicated in HA mode. -
Sealed store — Immutable, compressed columnar blocks on disk. The maintenance scheduler periodically compacts mutable data into sealed blocks using Gorilla, Delta-of-Delta, Simple-8b, and Dictionary codecs. Block-level min/max/sum statistics enable zero-decompression aggregation when an entire block falls within a single time bucket.
|
Since v26.11.1. Integer columns ( |
Data is distributed across N shards (default: arcadedb.asyncWorkerThreads, the available cores minus 1) for parallel writes and reads.
Each shard maintains its own mutable bucket and sealed store.
Queries merge results across shards using a min-heap priority queue sorted by timestamp.
See Also
-
Time Series Tutorial — Step-by-step hands-on guide to time series in ArcadeDB
-
Realtime Analytics — Use case for streaming ingestion and real-time dashboards