Ingest a PI, IP.21 or OPC export¶
This guide turns the CSV files a historian exports into archives. It covers local time, day-first dates, digital states in the value column, quality codes, interpolated exports, regional CSV formats, and wide and long files. Each step runs the real command on a small export.
The recipe¶
- Write a metadata template with
tsdive ingest <export> --init-meta META.json. - Fill in the template: identity, name, unit, engineering range, sample rate, quality codes.
- Ingest with the flags your export needs:
--tz,--dayfirstor--timestamp-format, and the column names. - Profile the archive and read the headline.
A PI export in local time¶
A PI DataLink export often has local timestamps in the regional date
order, and writes digital states such as I/O Timeout into the value
column. It has no quality column:
Timestamp,TI3101.PV
28/10/2026 06:00:00,181.2
28/10/2026 06:05:00,181.4
28/10/2026 06:10:00,I/O Timeout
28/10/2026 06:15:00,I/O Timeout
28/10/2026 06:20:00,181.9
28/10/2026 06:25:00,182.1
Write the template, naming the columns the export uses:
$ tsdive ingest TI3101.csv --init-meta TI3101.json \
--timestamp-col Timestamp --value-col TI3101.PV
wrote TI3101.json
The columns comment in the template says what ingest found: no
quality column. quality_codes lists I/O Timeout, the one string in
the value column, with no severity. Fill in the tag, and give the state
the severity BAD so ingest reads it as a BAD sample:
{
"identity": {"source_id": "pi-north", "point_id": "TI3101.PV"},
"name": "TI-3101 reactor outlet temperature",
"unit_raw": "degC",
"eng_range_zero": 0.0,
"eng_range_span": 250.0,
"sample_rate_s": 300,
"quality_codes": {"I/O Timeout": "BAD"}
}
With no quality column, --assume-quality GOOD states the quality of
every other sample. The archive records that the quality was assumed,
and every profile of it says so. --tz and --dayfirst place the local
dates in UTC:
$ tsdive ingest TI3101.csv --out TI3101.parquet --meta TI3101.json \
--timestamp-col Timestamp --value-col TI3101.PV \
--tz Europe/Berlin --dayfirst --assume-quality GOOD
wrote TI3101.parquet
tag pi-north:TI3101.PV 6 rows quality assumed GOOD
tz Europe/Berlin -> UTC
warning: quality assumed GOOD for every sample; the archive records this and every profile of it says so
$ tsdive profile TI3101.parquet
pi-north:TI3101.PV TI-3101 reactor outlet temperature
coverage 1.000 GOOD 4/6 censored no gaps 0
window 2026-10-28 05:00:00Z -> 05:25:00Z (25 min)
contract TIME_WEIGHTED RECORDED NONE stepped no digest d59d433d9c62
units degC -> degrees Celsius
Coverage
coverage 1.000 valid 0.667 gaps 0 data-loss gaps 0 longest n/a
Quality
GOOD 4 UNCERTAIN 0 BAD 2
quality ASSUMED at ingest: a caller's declaration, not observed
Range
clipped 0.0000 censored no
[10 more lines not shown]
GOOD 4/6 and BAD 2: the two I/O Timeout rows are BAD and their
values are nulled, so they stay out of the statistics. 06:00 in Berlin on
28 October 2026 is 05:00 UTC, because the clocks went back on 25 October.
An IP.21 export with its own date format¶
IP.21 and many SCADA systems write dates such as 28-Oct-2026 06:00:00.
--timestamp-format takes a
strptime format
for them:
TS,VALUE,STATUS
28-Oct-2026 06:00:00,3.52,Good
28-Oct-2026 06:01:00,3.55,Good
28-Oct-2026 06:02:00,3.51,Good
This export holds values the historian interpolated every minute, not
the values it stored. Say so with retrieval_mode, so every report of
the archive states it:
{
"identity": {"source_id": "ip21", "point_id": "PI4402.PV"},
"name": "PI-4402 reactor pressure",
"unit_raw": "bar",
"sample_rate_s": 60,
"retrieval_mode": "INTERPOLATED"
}
$ tsdive ingest PI4402.csv --out PI4402.parquet --meta PI4402.json \
--timestamp-col TS --value-col VALUE --quality-col STATUS \
--tz Europe/Berlin --timestamp-format "%d-%b-%Y %H:%M:%S"
wrote PI4402.parquet
tag ip21:PI4402.PV 3 rows quality from column STATUS
tz Europe/Berlin -> UTC
$ tsdive profile PI4402.parquet
ip21:PI4402.PV PI-4402 reactor pressure
coverage 1.000 GOOD 3/3 censored unknown gaps 0
window 2026-10-28 05:00:00Z -> 05:02:00Z (2 min)
contract TIME_WEIGHTED INTERPOLATED NONE stepped no digest 021659760bd6
[20 more lines not shown]
The contract line reads INTERPOLATED, and the digest changed with it.
A CSV saved by a German Excel¶
Excel in a German locale separates columns with ; and writes , as
the decimal mark:
Zeitstempel;Wert;Status
28.10.2026 06:00:00;181,2;Good
28.10.2026 06:05:00;181,4;Good
28.10.2026 06:10:00;181,9;Good
{
"identity": {"source_id": "pi-north", "point_id": "TI3102.PV"},
"name": "TI-3102 reactor inlet temperature",
"unit_raw": "degC",
"sample_rate_s": 300
}
--sep and --decimal state the format:
$ tsdive ingest TI3102.csv --out TI3102.parquet --meta TI3102.json \
--sep ";" --decimal "," --timestamp-col Zeitstempel --value-col Wert \
--quality-col Status --tz Europe/Berlin
wrote TI3102.parquet
tag pi-north:TI3102.PV 3 rows quality from column Status
tz Europe/Berlin -> UTC
Without --sep the header reads as one column, and ingest raises
SchemaError naming --sep. If the file holds a character such as °
in the Windows code page, pass --encoding cp1252 too.
An OPC export with byte quality codes¶
Classic OPC DA writes quality as a byte: 192 Good, 64 Uncertain, 0 Bad, with sub-status values around them. Read as OPC UA codes they mean something else, so declare them:
SourceTimestamp,Value,StatusCode
2026-10-28T05:00:00Z,12.40,192
2026-10-28T05:01:00Z,12.38,192
2026-10-28T05:02:00Z,12.41,216
2026-10-28T05:03:00Z,0.00,0
2026-10-28T05:04:00Z,12.39,192
2026-10-28T05:05:00Z,12.44,64
$ tsdive ingest FIC501.csv --init-meta FIC501.json \
--timestamp-col SourceTimestamp --value-col Value --quality-col StatusCode
wrote FIC501.json
The template lists every code the file holds, each waiting for a
severity. 216 is a Good sub-status (local override) in OPC DA:
{
"identity": {"source_id": "opc-line5", "point_id": "FIC501.PV"},
"name": "FIC-501 feed flow",
"unit_raw": "m3/h",
"sample_rate_s": 60,
"quality_codes": {"192": "GOOD", "216": "GOOD", "64": "UNCERTAIN", "0": "BAD"}
}
The timestamps carry Z, so no --tz is needed:
$ tsdive ingest FIC501.csv --out FIC501.parquet --meta FIC501.json \
--timestamp-col SourceTimestamp --value-col Value --quality-col StatusCode
wrote FIC501.parquet
tag opc-line5:FIC501.PV 6 rows quality from column StatusCode
$ tsdive profile FIC501.parquet
opc-line5:FIC501.PV FIC-501 feed flow
coverage 1.000 GOOD 4/6 censored unknown gaps 0
window 2026-10-28 05:00:00Z -> 05:05:00Z (5 min)
contract TIME_WEIGHTED RECORDED NONE stepped no digest d59d433d9c62
units m3/h -> cubic meters per hour
Coverage
coverage 1.000 valid 0.667 gaps 0 data-loss gaps 0 longest n/a
Quality
GOOD 4 UNCERTAIN 1 BAD 1
Range
[11 more lines not shown]
The BAD sample at 05:03 carries a value of 0.00 that is not a flow; it stays out of every statistic.
A wide export¶
A wide export has one column per tag, and its quality columns share a
suffix such as _q. Two commands handle it, as the
usage walkthrough shows: --wide --init-meta
DIR writes one template per tag, and --wide --out DIR --meta-dir DIR
writes one archive per tag.
A long export¶
A long export has one row per tag and timestamp, and a column that names the tag of each row:
Tag,Timestamp,Value,Status
FI102.PV,2026-10-28T05:00:00Z,40.06,Good
TI103.PV,2026-10-28T05:00:00Z,181.2,Good
FI102.PV,2026-10-28T05:01:00Z,40.11,Good
TI103.PV,2026-10-28T05:01:00Z,181.4,Good
FI102.PV,2026-10-28T05:02:00Z,40.02,Good
TI103.PV,2026-10-28T05:02:00Z,181.3,Good
A single-tag ingest of this file raises SchemaError, because two rows
share each timestamp and Tag splits them into two series. A file whose
tags sample at offset instants, such as :00 and :30, raises it too,
because consecutive rows move between the two series. --tag-col
names the tag column. With --init-meta DIR ingest writes one template
per tag:
$ tsdive ingest plant.csv --tag-col Tag --timestamp-col Timestamp \
--value-col Value --quality-col Status --init-meta meta --source-id plant1
wrote meta/FI102.PV.json
wrote meta/TI103.PV.json
{
"identity": {
"source_id": "plant1",
"point_id": "FI102.PV"
},
"name": "FI102.PV",
"unit_raw": null,
"unit_canonical": null,
"eng_range_zero": null,
"eng_range_span": null,
"sample_rate_s": null,
"retrieval_mode": null,
"asset": null,
"loop_id": null,
"role": null,
"quality_codes": {
"Good": "GOOD"
},
"quality_assumed": null
}
Fill in the unit, range and sample rate of each template. With
--out DIR --meta-dir DIR ingest writes one archive per tag, named by
point_id:
$ tsdive ingest plant.csv --tag-col Tag --timestamp-col Timestamp \
--value-col Value --quality-col Status --out archive --meta-dir meta
wrote archive/FI102.PV.parquet
wrote archive/TI103.PV.parquet
Read the result in a script¶
With --json ingest prints one object in place of the wrote lines.
form names the export shape, and archives holds one entry per
archive written: path, tag, identity, row count, first and last
timestamp, and quality_source (column or assumed). Under
--init-meta the object lists templates instead. A refused ingest
prints the refusal object and exits 3.
$ tsdive ingest plant.csv --tag-col Tag --timestamp-col Timestamp \
--value-col Value --quality-col Status --out archive --meta-dir meta \
--overwrite --json
{
"result_kind": "ingest",
"tsdive_version": "0.10.0",
"form": "long",
"archives": [
{
"path": "archive/FI102.PV.parquet",
"tag": "plant1:FI102.PV",
"identity": {
"source_id": "plant1",
"point_id": "FI102.PV"
},
"rows": 3,
"first": "2026-10-28T05:00:00+00:00",
"last": "2026-10-28T05:02:00+00:00",
"quality_source": "column",
"assumed_quality": null
},
{
"path": "archive/TI103.PV.parquet",
"tag": "plant1:TI103.PV",
"identity": {
"source_id": "plant1",
"point_id": "TI103.PV"
},
"rows": 3,
"first": "2026-10-28T05:00:00+00:00",
"last": "2026-10-28T05:02:00+00:00",
"quality_source": "column",
"assumed_quality": null
}
]
}
When ingest refuses¶
| message says | do |
|---|---|
the header reads as one column |
pass --sep with the separator the message names |
is not utf-8 |
pass --encoding with the encoding the export was written in |
is naive (no UTC offset) |
pass --tz with the zone the export was written in |
reads as day 02 of month 03 or as month 02 |
pass --dayfirst, or --timestamp-format |
repeats when the clocks go back |
export the stretch around the clock change with UTC offsets |
does not exist in |
export the stretch around the clock change with UTC offsets |
no quality column |
name it with --quality-col, or pass --assume-quality |
unknown key |
fix the metadata key it names; the message suggests the right one |
not numeric and no digital state |
give each string it lists a severity in quality_codes |
the export holds several tags |
pass --tag-col with the column the message names |
precedes the row before it |
fix the row order in the export; tsdive does not sort it |
The errors page lists every schema refusal.