Skip to content

Ingest a PI, IP.21 or OPC export

This guide turns the CSV files a historian exports into archives. It covers local time, day-first dates, digital states in the value column, quality codes, interpolated exports, regional CSV formats, and wide and long files. Each step runs the real command on a small export.

The recipe

  1. Write a metadata template with tsdive ingest <export> --init-meta META.json.
  2. Fill in the template: identity, name, unit, engineering range, sample rate, quality codes.
  3. Ingest with the flags your export needs: --tz, --dayfirst or --timestamp-format, and the column names.
  4. Profile the archive and read the headline.

A PI export in local time

A PI DataLink export often has local timestamps in the regional date order, and writes digital states such as I/O Timeout into the value column. It has no quality column:

TI3101.csv
Timestamp,TI3101.PV
28/10/2026 06:00:00,181.2
28/10/2026 06:05:00,181.4
28/10/2026 06:10:00,I/O Timeout
28/10/2026 06:15:00,I/O Timeout
28/10/2026 06:20:00,181.9
28/10/2026 06:25:00,182.1

Write the template, naming the columns the export uses:

$ tsdive ingest TI3101.csv --init-meta TI3101.json \
    --timestamp-col Timestamp --value-col TI3101.PV
wrote     TI3101.json

The columns comment in the template says what ingest found: no quality column. quality_codes lists I/O Timeout, the one string in the value column, with no severity. Fill in the tag, and give the state the severity BAD so ingest reads it as a BAD sample:

TI3101.json
{
  "identity": {"source_id": "pi-north", "point_id": "TI3101.PV"},
  "name": "TI-3101 reactor outlet temperature",
  "unit_raw": "degC",
  "eng_range_zero": 0.0,
  "eng_range_span": 250.0,
  "sample_rate_s": 300,
  "quality_codes": {"I/O Timeout": "BAD"}
}

With no quality column, --assume-quality GOOD states the quality of every other sample. The archive records that the quality was assumed, and every profile of it says so. --tz and --dayfirst place the local dates in UTC:

$ tsdive ingest TI3101.csv --out TI3101.parquet --meta TI3101.json \
    --timestamp-col Timestamp --value-col TI3101.PV \
    --tz Europe/Berlin --dayfirst --assume-quality GOOD
wrote     TI3101.parquet
tag       pi-north:TI3101.PV   6 rows   quality assumed GOOD
tz        Europe/Berlin -> UTC
warning: quality assumed GOOD for every sample; the archive records this and every profile of it says so
$ tsdive profile TI3101.parquet
pi-north:TI3101.PV  TI-3101 reactor outlet temperature
coverage 1.000   GOOD 4/6   censored no   gaps 0

window    2026-10-28 05:00:00Z -> 05:25:00Z  (25 min)
contract  TIME_WEIGHTED  RECORDED  NONE  stepped no  digest d59d433d9c62
units     degC -> degrees Celsius

Coverage
  coverage 1.000   valid 0.667   gaps 0   data-loss gaps 0   longest n/a

Quality
  GOOD 4   UNCERTAIN 0   BAD 2
  quality ASSUMED at ingest: a caller's declaration, not observed

Range
  clipped 0.0000   censored no
[10 more lines not shown]

GOOD 4/6 and BAD 2: the two I/O Timeout rows are BAD and their values are nulled, so they stay out of the statistics. 06:00 in Berlin on 28 October 2026 is 05:00 UTC, because the clocks went back on 25 October.

An IP.21 export with its own date format

IP.21 and many SCADA systems write dates such as 28-Oct-2026 06:00:00. --timestamp-format takes a strptime format for them:

PI4402.csv
TS,VALUE,STATUS
28-Oct-2026 06:00:00,3.52,Good
28-Oct-2026 06:01:00,3.55,Good
28-Oct-2026 06:02:00,3.51,Good

This export holds values the historian interpolated every minute, not the values it stored. Say so with retrieval_mode, so every report of the archive states it:

PI4402.json
{
  "identity": {"source_id": "ip21", "point_id": "PI4402.PV"},
  "name": "PI-4402 reactor pressure",
  "unit_raw": "bar",
  "sample_rate_s": 60,
  "retrieval_mode": "INTERPOLATED"
}
$ tsdive ingest PI4402.csv --out PI4402.parquet --meta PI4402.json \
    --timestamp-col TS --value-col VALUE --quality-col STATUS \
    --tz Europe/Berlin --timestamp-format "%d-%b-%Y %H:%M:%S"
wrote     PI4402.parquet
tag       ip21:PI4402.PV   3 rows   quality from column STATUS
tz        Europe/Berlin -> UTC
$ tsdive profile PI4402.parquet
ip21:PI4402.PV  PI-4402 reactor pressure
coverage 1.000   GOOD 3/3   censored unknown   gaps 0

window    2026-10-28 05:00:00Z -> 05:02:00Z  (2 min)
contract  TIME_WEIGHTED  INTERPOLATED  NONE  stepped no  digest 021659760bd6
[20 more lines not shown]

The contract line reads INTERPOLATED, and the digest changed with it.

A CSV saved by a German Excel

Excel in a German locale separates columns with ; and writes , as the decimal mark:

TI3102.csv
Zeitstempel;Wert;Status
28.10.2026 06:00:00;181,2;Good
28.10.2026 06:05:00;181,4;Good
28.10.2026 06:10:00;181,9;Good
TI3102.json
{
  "identity": {"source_id": "pi-north", "point_id": "TI3102.PV"},
  "name": "TI-3102 reactor inlet temperature",
  "unit_raw": "degC",
  "sample_rate_s": 300
}

--sep and --decimal state the format:

$ tsdive ingest TI3102.csv --out TI3102.parquet --meta TI3102.json \
    --sep ";" --decimal "," --timestamp-col Zeitstempel --value-col Wert \
    --quality-col Status --tz Europe/Berlin
wrote     TI3102.parquet
tag       pi-north:TI3102.PV   3 rows   quality from column Status
tz        Europe/Berlin -> UTC

Without --sep the header reads as one column, and ingest raises SchemaError naming --sep. If the file holds a character such as ° in the Windows code page, pass --encoding cp1252 too.

An OPC export with byte quality codes

Classic OPC DA writes quality as a byte: 192 Good, 64 Uncertain, 0 Bad, with sub-status values around them. Read as OPC UA codes they mean something else, so declare them:

FIC501.csv
SourceTimestamp,Value,StatusCode
2026-10-28T05:00:00Z,12.40,192
2026-10-28T05:01:00Z,12.38,192
2026-10-28T05:02:00Z,12.41,216
2026-10-28T05:03:00Z,0.00,0
2026-10-28T05:04:00Z,12.39,192
2026-10-28T05:05:00Z,12.44,64
$ tsdive ingest FIC501.csv --init-meta FIC501.json \
    --timestamp-col SourceTimestamp --value-col Value --quality-col StatusCode
wrote     FIC501.json

The template lists every code the file holds, each waiting for a severity. 216 is a Good sub-status (local override) in OPC DA:

FIC501.json
{
  "identity": {"source_id": "opc-line5", "point_id": "FIC501.PV"},
  "name": "FIC-501 feed flow",
  "unit_raw": "m3/h",
  "sample_rate_s": 60,
  "quality_codes": {"192": "GOOD", "216": "GOOD", "64": "UNCERTAIN", "0": "BAD"}
}

The timestamps carry Z, so no --tz is needed:

$ tsdive ingest FIC501.csv --out FIC501.parquet --meta FIC501.json \
    --timestamp-col SourceTimestamp --value-col Value --quality-col StatusCode
wrote     FIC501.parquet
tag       opc-line5:FIC501.PV   6 rows   quality from column StatusCode
$ tsdive profile FIC501.parquet
opc-line5:FIC501.PV  FIC-501 feed flow
coverage 1.000   GOOD 4/6   censored unknown   gaps 0

window    2026-10-28 05:00:00Z -> 05:05:00Z  (5 min)
contract  TIME_WEIGHTED  RECORDED  NONE  stepped no  digest d59d433d9c62
units     m3/h -> cubic meters per hour

Coverage
  coverage 1.000   valid 0.667   gaps 0   data-loss gaps 0   longest n/a

Quality
  GOOD 4   UNCERTAIN 1   BAD 1

Range
[11 more lines not shown]

The BAD sample at 05:03 carries a value of 0.00 that is not a flow; it stays out of every statistic.

A wide export

A wide export has one column per tag, and its quality columns share a suffix such as _q. Two commands handle it, as the usage walkthrough shows: --wide --init-meta DIR writes one template per tag, and --wide --out DIR --meta-dir DIR writes one archive per tag.

A long export

A long export has one row per tag and timestamp, and a column that names the tag of each row:

plant.csv
Tag,Timestamp,Value,Status
FI102.PV,2026-10-28T05:00:00Z,40.06,Good
TI103.PV,2026-10-28T05:00:00Z,181.2,Good
FI102.PV,2026-10-28T05:01:00Z,40.11,Good
TI103.PV,2026-10-28T05:01:00Z,181.4,Good
FI102.PV,2026-10-28T05:02:00Z,40.02,Good
TI103.PV,2026-10-28T05:02:00Z,181.3,Good

A single-tag ingest of this file raises SchemaError, because two rows share each timestamp and Tag splits them into two series. A file whose tags sample at offset instants, such as :00 and :30, raises it too, because consecutive rows move between the two series. --tag-col names the tag column. With --init-meta DIR ingest writes one template per tag:

$ tsdive ingest plant.csv --tag-col Tag --timestamp-col Timestamp \
    --value-col Value --quality-col Status --init-meta meta --source-id plant1
wrote     meta/FI102.PV.json
wrote     meta/TI103.PV.json
meta/FI102.PV.json
{
  "identity": {
    "source_id": "plant1",
    "point_id": "FI102.PV"
  },
  "name": "FI102.PV",
  "unit_raw": null,
  "unit_canonical": null,
  "eng_range_zero": null,
  "eng_range_span": null,
  "sample_rate_s": null,
  "retrieval_mode": null,
  "asset": null,
  "loop_id": null,
  "role": null,
  "quality_codes": {
    "Good": "GOOD"
  },
  "quality_assumed": null
}

Fill in the unit, range and sample rate of each template. With --out DIR --meta-dir DIR ingest writes one archive per tag, named by point_id:

$ tsdive ingest plant.csv --tag-col Tag --timestamp-col Timestamp \
    --value-col Value --quality-col Status --out archive --meta-dir meta
wrote     archive/FI102.PV.parquet
wrote     archive/TI103.PV.parquet

Read the result in a script

With --json ingest prints one object in place of the wrote lines. form names the export shape, and archives holds one entry per archive written: path, tag, identity, row count, first and last timestamp, and quality_source (column or assumed). Under --init-meta the object lists templates instead. A refused ingest prints the refusal object and exits 3.

$ tsdive ingest plant.csv --tag-col Tag --timestamp-col Timestamp \
    --value-col Value --quality-col Status --out archive --meta-dir meta \
    --overwrite --json
{
  "result_kind": "ingest",
  "tsdive_version": "0.10.0",
  "form": "long",
  "archives": [
    {
      "path": "archive/FI102.PV.parquet",
      "tag": "plant1:FI102.PV",
      "identity": {
        "source_id": "plant1",
        "point_id": "FI102.PV"
      },
      "rows": 3,
      "first": "2026-10-28T05:00:00+00:00",
      "last": "2026-10-28T05:02:00+00:00",
      "quality_source": "column",
      "assumed_quality": null
    },
    {
      "path": "archive/TI103.PV.parquet",
      "tag": "plant1:TI103.PV",
      "identity": {
        "source_id": "plant1",
        "point_id": "TI103.PV"
      },
      "rows": 3,
      "first": "2026-10-28T05:00:00+00:00",
      "last": "2026-10-28T05:02:00+00:00",
      "quality_source": "column",
      "assumed_quality": null
    }
  ]
}

When ingest refuses

message says do
the header reads as one column pass --sep with the separator the message names
is not utf-8 pass --encoding with the encoding the export was written in
is naive (no UTC offset) pass --tz with the zone the export was written in
reads as day 02 of month 03 or as month 02 pass --dayfirst, or --timestamp-format
repeats when the clocks go back export the stretch around the clock change with UTC offsets
does not exist in export the stretch around the clock change with UTC offsets
no quality column name it with --quality-col, or pass --assume-quality
unknown key fix the metadata key it names; the message suggests the right one
not numeric and no digital state give each string it lists a severity in quality_codes
the export holds several tags pass --tag-col with the column the message names
precedes the row before it fix the row order in the export; tsdive does not sort it

The errors page lists every schema refusal.