Skip to content

Getting started

This tutorial takes you from an empty Python environment to a profile of your own historian export in about 15 minutes. Every command on this page ran when the page was built, and the block under it shows what it printed.

Install

tsdive needs Python 3.12 or newer. Check the version you have:

python --version

If it prints 3.11 or older, let uv fetch Python 3.12 into a project folder. It needs no administrator rights and leaves the system Python alone:

uv venv --python 3.12
.venv\Scripts\activate
uv pip install https://github.com/NorthernLightx/tsdive/releases/download/v0.10.0/tsdive-0.10.0-py3-none-any.whl

On Linux and macOS the second line is source .venv/bin/activate. With Python 3.12 already installed, pip install takes the same wheel URL. With git installed, pip install "git+https://github.com/NorthernLightx/tsdive.git" builds the current main branch instead. The package is not on PyPI.

Check the install:

$ tsdive --version
tsdive 0.10.0

Get the demo data

tsdive demo writes five small archives into a new folder. They are synthetic, so every number on this site comes out the same on your machine:

$ tsdive demo
wrote     tsdive-demo/demo/fic101_demo.parquet   562 samples
          tsdive-demo/demo/tic101_demo.parquet   562 samples
          tsdive-demo/switchback_demo/ti201.parquet   4320 samples
          tsdive-demo/switchback_demo/fi200.parquet   4320 samples
          tsdive-demo/switchback_demo/tt001.parquet   4320 samples
next      tsdive profile tsdive-demo/demo/fic101_demo.parquet

The two archives under demo/ hold ten hours of one flow loop, FIC-101 in m3/h, and one temperature, TIC-101 in degC, one sample a minute. The flow has three faults built in: a 40-minute collection outage, half an hour pinned at the top of its 0 to 100 m3/h range, and one sample whose quality code the site never defined. The three archives under switchback_demo/ belong to the switchback trial how-to.

An archive is one parquet file per tag: the samples, their raw quality codes, and the tag's metadata. The last line of the output is the command to try next.

Profile a window

tsdive profile reads one tag over a window and reports what the historian did to it:

$ tsdive profile tsdive-demo/demo/fic101_demo.parquet
demo:FIC101.PV  FIC-101 flow
coverage 0.933   GOOD 561/562   censored yes   gaps 1

window    2024-03-30 20:00:00Z -> 2024-03-31 06:00:00Z  (10 h)
contract  TIME_WEIGHTED  RECORDED  NONE  stepped no  digest d59d433d9c62
units     m3/h -> cubic meters per hour

Coverage
  coverage 0.933   valid 0.998   gaps 1   data-loss gaps 1   longest 40 min
  2024-03-30 23:00:00Z -> 23:40:00Z   40 min   unknown (no rule matched)

Quality
  GOOD 561   UNCERTAIN 1   BAD 0
  unmapped codes, treated UNCERTAIN: SENSOR DRIFT

Range
  clipped 0.0516   censored yes

Timestamps
  audited 562   duplicates 0   non-monotonic 0

Values  GOOD n=561
  min 60.86   p05 61.20   median 62.32   p95 100.0   max 100.0
  mean 63.95 (time-weighted)   std 8.399   mad 0.4017
  distinct 533   stall 0 s   changes/h 53.20
  constant run 2024-03-31 02:01:00Z -> 02:29:00Z   28 min   n=29
  interval 60 s (p05 60 s, p95 60 s)   declared 60 s

With no --window, the window is the archive's own extent, 20:00 to 06:00 UTC. Read the top of the report line by line:

  • demo:FIC101.PV FIC-101 flow is the tag's identity, the historian (demo) and the point (FIC101.PV), then its display name.
  • coverage 0.933 is the share of the 10-hour window that is not lost to a data-loss gap. The 40-minute outage takes 40 of 600 minutes, so coverage is 1 - 40/600.
  • GOOD 561/562 counts the samples whose quality code reads GOOD. The one other sample carries the undefined code SENSOR DRIFT, which tsdive counts as UNCERTAIN.
  • censored yes says some samples sit at the end of the engineering range, where the transmitter stops measuring. Those samples say only that the flow was at least 100 m3/h. See clipping and censoring.
  • gaps 1 counts every hole in the timestamps longer than the gap threshold, here the outage.
  • window, contract and units state what was read: the window in UTC, the sampling contract the statistics assume, and the unit resolved from the metadata.

The sections under it give the evidence behind each number. Reading the profile explains every line.

Screen a later window

tsdive screen asks whether a window behaves like a baseline. The baseline is a stretch of history you trust. Here it is the five hours from 20:00 to 01:00, and the window is the five hours after it:

$ tsdive screen tsdive-demo/demo/fic101_demo.parquet \
    --baseline 2024-03-30T20:00:00Z/2024-03-31T01:00:00Z \
    --window 2024-03-31T01:00:00Z/2024-03-31T06:00:00Z
demo:FIC101.PV  flagged 29 of 300 (9.7%)

baseline  2024-03-30 20:00:00Z -> 2024-03-31 01:00:00Z   GOOD 261   censored no
window    2024-03-31 01:00:00Z -> 06:00:00Z
method    MAD   center 62.32   scale 0.5488   k 3.0   limits [60.67, 63.97]
caveat    provisional: one baseline for every regime in the window

Flagged
  2024-03-31 02:01:00Z -> 02:29:00Z   29 samples

The baseline's median is 62.32 m3/h. Its spread, 1.4826 times the median absolute deviation, is 0.5488 m3/h, and that is the sigma the screen works in. A sample more than 3 sigma from the median is flagged, so the limits are 60.67 and 63.97 m3/h. The 29 flagged samples are the half hour the transmitter sat at 100 m3/h. Baselines and the MAD screen explains the method and the provisional caveat.

A baseline has to be clean. Move it over the saturated half hour and tsdive stops with a typed error instead of computing limits from it:

$ tsdive screen tsdive-demo/demo/fic101_demo.parquet \
    --baseline 2024-03-31T00:00:00Z/2024-03-31T03:00:00Z \
    --window 2024-03-31T03:00:00Z/2024-03-31T06:00:00Z
[InsufficientQuality] tag demo:FIC101.PV: window is censored (clipped fraction 0.160); a clipped window may never serve as a baseline

This is a refusal: the question has no answer on this data, and the message names the check and the reason. The exit status is 3. Pick a baseline window without the clipped stretch. Choose a clean baseline window shows how to find one.

Now your own data

Say your historian exports one tag to CSV, with local timestamps in day-first order and its own quality words:

FI2201.csv
Timestamp,Value,Quality
02/03/2026 08:00,41.30,Good
02/03/2026 08:05,41.17,Good
02/03/2026 08:10,41.52,Good
02/03/2026 08:15,42.32,Good
02/03/2026 08:20,42.06,Good
02/03/2026 08:25,42.22,Good
02/03/2026 08:30,42.80,Good
02/03/2026 08:35,42.28,Questionable
02/03/2026 08:40,42.16,Good
02/03/2026 08:45,42.47,Good
02/03/2026 08:50,41.70,Good
02/03/2026 08:55,41.37,Good
02/03/2026 09:00,41.51,Good
02/03/2026 09:05,40.64,Good
02/03/2026 09:30,39.83,Bad
02/03/2026 09:35,39.30,Good
02/03/2026 09:40,39.36,Good
02/03/2026 09:45,40.01,Good
02/03/2026 09:50,39.74,Good
02/03/2026 09:55,40.04,Good
02/03/2026 10:00,40.88,Good

tsdive reads an archive, not a CSV, so the first step is tsdive ingest. An archive carries the tag's metadata, and ingest writes a template for it:

$ tsdive ingest FI2201.csv --init-meta FI2201.json \
    --timestamp-col Timestamp --value-col Value --quality-col Quality
wrote     FI2201.json
FI2201.json
{
  "_comments": {
    "identity": "source_id names the historian or collector, point_id the tag in it; both required, never renamed",
    "name": "display name; required",
    "unit_raw": "unit exactly as the historian writes it; null leaves the archive without a unit",
    "unit_canonical": "leave null; tsdive resolves unit_raw itself",
    "eng_range_zero": "bottom of the engineering range; with eng_range_span it enables the clipping check",
    "eng_range_span": "width of the engineering range, greater than 0",
    "sample_rate_s": "declared scan rate in seconds; compare and mspc align on it",
    "retrieval_mode": "RECORDED or INTERPOLATED, how the export retrieved its samples; null reads as RECORDED",
    "asset": "unit or equipment the tag belongs to",
    "loop_id": "control loop id",
    "role": "PV, SP, OP or MODE; MODE for a tag whose values are string states",
    "quality_codes": "each raw quality code, and each string in a numeric value column, mapped to GOOD, UNCERTAIN or BAD; a null entry raises SchemaError until it names a severity",
    "quality_assumed": "leave null; ingest sets it when the quality is assumed",
    "columns": "timestamp 'Timestamp', value 'Value', quality 'Quality'; the export has Timestamp, Value, Quality"
  },
  "identity": {
    "source_id": null,
    "point_id": null
  },
  "name": null,
  "unit_raw": null,
  "unit_canonical": null,
  "eng_range_zero": null,
  "eng_range_span": null,
  "sample_rate_s": null,
  "retrieval_mode": null,
  "asset": null,
  "loop_id": null,
  "role": null,
  "quality_codes": {
    "Bad": "BAD",
    "Good": "GOOD",
    "Questionable": null
  },
  "quality_assumed": null
}

The _comments block explains each key and ingest ignores it. The quality codes come from the file: Good and Bad are mapped already, and Questionable waits for you to name its severity. Fill in what you know about the tag and delete what you do not:

FI2201.json
{
  "identity": {"source_id": "plant1", "point_id": "FI2201.PV"},
  "name": "FI-2201 cooling water flow",
  "unit_raw": "m3/h",
  "eng_range_zero": 0.0,
  "eng_range_span": 80.0,
  "sample_rate_s": 300,
  "role": "PV",
  "quality_codes": {"Good": "GOOD", "Questionable": "UNCERTAIN", "Bad": "BAD"}
}

The engineering range turns on the clipping check, and the sample rate of 300 s tells tsdive that a 5-minute spacing is normal. The timestamps carry no offset, so --tz states the zone they were written in:

$ tsdive ingest FI2201.csv --out FI2201.parquet --meta FI2201.json \
    --timestamp-col Timestamp --value-col Value --quality-col Quality \
    --tz Europe/Berlin
[SchemaError] Timestamp: '02/03/2026 08:00' reads as day 02 of month 03 or as month 02, day 03 of 2026; pass --dayfirst to read day first, or --timestamp-format with a strptime format such as '%d/%m/%Y %H:%M'

02/03/2026 is 2 March in Europe and 3 February in the US, and tsdive does not pick one. State the order with --dayfirst:

$ tsdive ingest FI2201.csv --out FI2201.parquet --meta FI2201.json \
    --timestamp-col Timestamp --value-col Value --quality-col Quality \
    --tz Europe/Berlin --dayfirst
wrote     FI2201.parquet
tag       plant1:FI2201.PV   21 rows   quality from column Quality
tz        Europe/Berlin -> UTC

The archive holds UTC timestamps. 08:00 in Berlin in March is 07:00 UTC. Profile it:

$ tsdive profile FI2201.parquet
plant1:FI2201.PV  FI-2201 cooling water flow
coverage 0.792   GOOD 19/21   censored no   gaps 1

window    2026-03-02 07:00:00Z -> 09:00:00Z  (2 h)
contract  TIME_WEIGHTED  RECORDED  NONE  stepped no  digest d59d433d9c62
units     m3/h -> cubic meters per hour

Coverage
  coverage 0.792   valid 0.905   gaps 1   data-loss gaps 1   longest 25 min
  2026-03-02 08:05:00Z -> 08:30:00Z   25 min   unknown (no rule matched)

Quality
  GOOD 19   UNCERTAIN 1   BAD 1

Range
  clipped 0.0000   censored no

Timestamps
  audited 21   duplicates 0   non-monotonic 0

Values  GOOD n=19
  min 39.30   p05 39.35   median 41.37   p95 42.50   max 42.80
  mean 41.15 (time-weighted)   std 1.078   mad 0.7900
  distinct 19   stall 0 s   changes/h 9.000
  constant run 2026-03-02 07:00:00Z -> 07:00:00Z   0 s   n=1
  interval 300 s (p05 300 s, p95 360 s)   declared 300 s

Coverage is 0.792, because the 25 minutes between 09:05 and 09:30 local time hold no sample, where a 5-minute tag should have four. One sample is UNCERTAIN and one BAD, so 19 of 21 are GOOD, and only those 19 feed the statistics. censored no is now a real answer: the range is declared and no sample reached 0 or 80 m3/h.

Next

  • Reading the output explains every line of profile, screen, spc, mspc, compare and switchback analyze.
  • The how-to guides cover real exports, baseline choice, many tags at once, Python, AI assistants and controller trials.
  • Errors and refusals lists every typed error and what to do about it.