Getting started¶
This tutorial takes you from an empty Python environment to a profile of your own historian export in about 15 minutes. Every command on this page ran when the page was built, and the block under it shows what it printed.
Install¶
tsdive needs Python 3.12 or newer. Check the version you have:
python --version
If it prints 3.11 or older, let uv fetch Python 3.12 into a project folder. It needs no administrator rights and leaves the system Python alone:
uv venv --python 3.12
.venv\Scripts\activate
uv pip install https://github.com/NorthernLightx/tsdive/releases/download/v0.10.0/tsdive-0.10.0-py3-none-any.whl
On Linux and macOS the second line is source .venv/bin/activate. With
Python 3.12 already installed, pip install takes the same wheel URL.
With git installed, pip install "git+https://github.com/NorthernLightx/tsdive.git"
builds the current main branch instead. The package is not on PyPI.
Check the install:
$ tsdive --version
tsdive 0.10.0
Get the demo data¶
tsdive demo writes five small archives into a new folder. They are
synthetic, so every number on this site comes out the same on your
machine:
$ tsdive demo
wrote tsdive-demo/demo/fic101_demo.parquet 562 samples
tsdive-demo/demo/tic101_demo.parquet 562 samples
tsdive-demo/switchback_demo/ti201.parquet 4320 samples
tsdive-demo/switchback_demo/fi200.parquet 4320 samples
tsdive-demo/switchback_demo/tt001.parquet 4320 samples
next tsdive profile tsdive-demo/demo/fic101_demo.parquet
The two archives under demo/ hold ten hours of one flow loop, FIC-101
in m3/h, and one temperature, TIC-101 in degC, one sample a minute. The
flow has three faults built in: a 40-minute collection outage, half an
hour pinned at the top of its 0 to 100 m3/h range, and one sample whose
quality code the site never defined. The three archives under
switchback_demo/ belong to the switchback trial how-to.
An archive is one parquet file per tag: the samples, their raw quality codes, and the tag's metadata. The last line of the output is the command to try next.
Profile a window¶
tsdive profile reads one tag over a window and reports what the
historian did to it:
$ tsdive profile tsdive-demo/demo/fic101_demo.parquet
demo:FIC101.PV FIC-101 flow
coverage 0.933 GOOD 561/562 censored yes gaps 1
window 2024-03-30 20:00:00Z -> 2024-03-31 06:00:00Z (10 h)
contract TIME_WEIGHTED RECORDED NONE stepped no digest d59d433d9c62
units m3/h -> cubic meters per hour
Coverage
coverage 0.933 valid 0.998 gaps 1 data-loss gaps 1 longest 40 min
2024-03-30 23:00:00Z -> 23:40:00Z 40 min unknown (no rule matched)
Quality
GOOD 561 UNCERTAIN 1 BAD 0
unmapped codes, treated UNCERTAIN: SENSOR DRIFT
Range
clipped 0.0516 censored yes
Timestamps
audited 562 duplicates 0 non-monotonic 0
Values GOOD n=561
min 60.86 p05 61.20 median 62.32 p95 100.0 max 100.0
mean 63.95 (time-weighted) std 8.399 mad 0.4017
distinct 533 stall 0 s changes/h 53.20
constant run 2024-03-31 02:01:00Z -> 02:29:00Z 28 min n=29
interval 60 s (p05 60 s, p95 60 s) declared 60 s
With no --window, the window is the archive's own extent, 20:00 to
06:00 UTC. Read the top of the report line by line:
demo:FIC101.PV FIC-101 flowis the tag's identity, the historian (demo) and the point (FIC101.PV), then its display name.coverage 0.933is the share of the 10-hour window that is not lost to a data-loss gap. The 40-minute outage takes 40 of 600 minutes, so coverage is 1 - 40/600.GOOD 561/562counts the samples whose quality code reads GOOD. The one other sample carries the undefined codeSENSOR DRIFT, which tsdive counts as UNCERTAIN.censored yessays some samples sit at the end of the engineering range, where the transmitter stops measuring. Those samples say only that the flow was at least 100 m3/h. See clipping and censoring.gaps 1counts every hole in the timestamps longer than the gap threshold, here the outage.window,contractandunitsstate what was read: the window in UTC, the sampling contract the statistics assume, and the unit resolved from the metadata.
The sections under it give the evidence behind each number. Reading the profile explains every line.
Screen a later window¶
tsdive screen asks whether a window behaves like a baseline. The
baseline is a stretch of history you
trust. Here it is the five hours from 20:00 to 01:00, and the window is
the five hours after it:
$ tsdive screen tsdive-demo/demo/fic101_demo.parquet \
--baseline 2024-03-30T20:00:00Z/2024-03-31T01:00:00Z \
--window 2024-03-31T01:00:00Z/2024-03-31T06:00:00Z
demo:FIC101.PV flagged 29 of 300 (9.7%)
baseline 2024-03-30 20:00:00Z -> 2024-03-31 01:00:00Z GOOD 261 censored no
window 2024-03-31 01:00:00Z -> 06:00:00Z
method MAD center 62.32 scale 0.5488 k 3.0 limits [60.67, 63.97]
caveat provisional: one baseline for every regime in the window
Flagged
2024-03-31 02:01:00Z -> 02:29:00Z 29 samples
The baseline's median is 62.32 m3/h. Its spread, 1.4826 times the
median absolute deviation, is 0.5488 m3/h, and that is the sigma the
screen works in. A sample more than 3 sigma from the median is flagged,
so the limits are 60.67 and 63.97 m3/h. The 29 flagged samples are the
half hour the transmitter sat at 100 m3/h.
Baselines and the MAD screen explains the
method and the provisional caveat.
A baseline has to be clean. Move it over the saturated half hour and tsdive stops with a typed error instead of computing limits from it:
$ tsdive screen tsdive-demo/demo/fic101_demo.parquet \
--baseline 2024-03-31T00:00:00Z/2024-03-31T03:00:00Z \
--window 2024-03-31T03:00:00Z/2024-03-31T06:00:00Z
[InsufficientQuality] tag demo:FIC101.PV: window is censored (clipped fraction 0.160); a clipped window may never serve as a baseline
This is a refusal: the question has no answer on this data, and the message names the check and the reason. The exit status is 3. Pick a baseline window without the clipped stretch. Choose a clean baseline window shows how to find one.
Now your own data¶
Say your historian exports one tag to CSV, with local timestamps in day-first order and its own quality words:
Timestamp,Value,Quality
02/03/2026 08:00,41.30,Good
02/03/2026 08:05,41.17,Good
02/03/2026 08:10,41.52,Good
02/03/2026 08:15,42.32,Good
02/03/2026 08:20,42.06,Good
02/03/2026 08:25,42.22,Good
02/03/2026 08:30,42.80,Good
02/03/2026 08:35,42.28,Questionable
02/03/2026 08:40,42.16,Good
02/03/2026 08:45,42.47,Good
02/03/2026 08:50,41.70,Good
02/03/2026 08:55,41.37,Good
02/03/2026 09:00,41.51,Good
02/03/2026 09:05,40.64,Good
02/03/2026 09:30,39.83,Bad
02/03/2026 09:35,39.30,Good
02/03/2026 09:40,39.36,Good
02/03/2026 09:45,40.01,Good
02/03/2026 09:50,39.74,Good
02/03/2026 09:55,40.04,Good
02/03/2026 10:00,40.88,Good
tsdive reads an archive, not a CSV, so the first step is
tsdive ingest. An archive carries the tag's
metadata, and ingest writes a template for it:
$ tsdive ingest FI2201.csv --init-meta FI2201.json \
--timestamp-col Timestamp --value-col Value --quality-col Quality
wrote FI2201.json
{
"_comments": {
"identity": "source_id names the historian or collector, point_id the tag in it; both required, never renamed",
"name": "display name; required",
"unit_raw": "unit exactly as the historian writes it; null leaves the archive without a unit",
"unit_canonical": "leave null; tsdive resolves unit_raw itself",
"eng_range_zero": "bottom of the engineering range; with eng_range_span it enables the clipping check",
"eng_range_span": "width of the engineering range, greater than 0",
"sample_rate_s": "declared scan rate in seconds; compare and mspc align on it",
"retrieval_mode": "RECORDED or INTERPOLATED, how the export retrieved its samples; null reads as RECORDED",
"asset": "unit or equipment the tag belongs to",
"loop_id": "control loop id",
"role": "PV, SP, OP or MODE; MODE for a tag whose values are string states",
"quality_codes": "each raw quality code, and each string in a numeric value column, mapped to GOOD, UNCERTAIN or BAD; a null entry raises SchemaError until it names a severity",
"quality_assumed": "leave null; ingest sets it when the quality is assumed",
"columns": "timestamp 'Timestamp', value 'Value', quality 'Quality'; the export has Timestamp, Value, Quality"
},
"identity": {
"source_id": null,
"point_id": null
},
"name": null,
"unit_raw": null,
"unit_canonical": null,
"eng_range_zero": null,
"eng_range_span": null,
"sample_rate_s": null,
"retrieval_mode": null,
"asset": null,
"loop_id": null,
"role": null,
"quality_codes": {
"Bad": "BAD",
"Good": "GOOD",
"Questionable": null
},
"quality_assumed": null
}
The _comments block explains each key and ingest ignores it. The
quality codes come from the file: Good and Bad are mapped already,
and Questionable waits for you to name its severity. Fill in what you
know about the tag and delete what you do not:
{
"identity": {"source_id": "plant1", "point_id": "FI2201.PV"},
"name": "FI-2201 cooling water flow",
"unit_raw": "m3/h",
"eng_range_zero": 0.0,
"eng_range_span": 80.0,
"sample_rate_s": 300,
"role": "PV",
"quality_codes": {"Good": "GOOD", "Questionable": "UNCERTAIN", "Bad": "BAD"}
}
The engineering range turns on the clipping check, and the sample rate
of 300 s tells tsdive that a 5-minute spacing is normal. The timestamps
carry no offset, so --tz states the zone they were written in:
$ tsdive ingest FI2201.csv --out FI2201.parquet --meta FI2201.json \
--timestamp-col Timestamp --value-col Value --quality-col Quality \
--tz Europe/Berlin
[SchemaError] Timestamp: '02/03/2026 08:00' reads as day 02 of month 03 or as month 02, day 03 of 2026; pass --dayfirst to read day first, or --timestamp-format with a strptime format such as '%d/%m/%Y %H:%M'
02/03/2026 is 2 March in Europe and 3 February in the US, and tsdive
does not pick one. State the order with --dayfirst:
$ tsdive ingest FI2201.csv --out FI2201.parquet --meta FI2201.json \
--timestamp-col Timestamp --value-col Value --quality-col Quality \
--tz Europe/Berlin --dayfirst
wrote FI2201.parquet
tag plant1:FI2201.PV 21 rows quality from column Quality
tz Europe/Berlin -> UTC
The archive holds UTC timestamps. 08:00 in Berlin in March is 07:00 UTC. Profile it:
$ tsdive profile FI2201.parquet
plant1:FI2201.PV FI-2201 cooling water flow
coverage 0.792 GOOD 19/21 censored no gaps 1
window 2026-03-02 07:00:00Z -> 09:00:00Z (2 h)
contract TIME_WEIGHTED RECORDED NONE stepped no digest d59d433d9c62
units m3/h -> cubic meters per hour
Coverage
coverage 0.792 valid 0.905 gaps 1 data-loss gaps 1 longest 25 min
2026-03-02 08:05:00Z -> 08:30:00Z 25 min unknown (no rule matched)
Quality
GOOD 19 UNCERTAIN 1 BAD 1
Range
clipped 0.0000 censored no
Timestamps
audited 21 duplicates 0 non-monotonic 0
Values GOOD n=19
min 39.30 p05 39.35 median 41.37 p95 42.50 max 42.80
mean 41.15 (time-weighted) std 1.078 mad 0.7900
distinct 19 stall 0 s changes/h 9.000
constant run 2026-03-02 07:00:00Z -> 07:00:00Z 0 s n=1
interval 300 s (p05 300 s, p95 360 s) declared 300 s
Coverage is 0.792, because the 25 minutes between 09:05 and 09:30 local
time hold no sample, where a 5-minute tag should have four. One sample
is UNCERTAIN and one BAD, so 19 of 21 are GOOD, and only those 19 feed
the statistics. censored no is now a real answer: the range is
declared and no sample reached 0 or 80 m3/h.
Next¶
- Reading the output explains every line of profile, screen, spc, mspc, compare and switchback analyze.
- The how-to guides cover real exports, baseline choice, many tags at once, Python, AI assistants and controller trials.
- Errors and refusals lists every typed error and what to do about it.