API reference / Reading and ingest
tsdive.ingest
¶
ingest(
source: str | Path,
*,
out: str | Path,
meta: TagMeta,
timestamp_col: str = "timestamp",
value_col: str = "value",
quality_col: str = "quality",
tz: str | None = None,
assume_quality: str | None = None,
overwrite: bool = False,
timestamp_format: str | None = None,
dayfirst: bool = False,
sep: str | None = None,
decimal: str | None = None,
encoding: str | None = None,
) -> Path
Turn a CSV or parquet export into a tsdive archive.
Only the three declared columns are carried over; everything else in
the export is left behind, so an ingested archive contains exactly
what its schema promises. A left-behind column that splits the rows
into overlapping series, each in time order, marks an export of
several tags. Such an export raises SchemaError when its
timestamps repeat or run backwards, or when more than half of its
consecutive rows move to another series, as tags on offset clocks do.
ingest_long writes one archive per tag from
it. The offset-clock rule skips a column of floats, which holds
measurements. Timestamps that run backwards
otherwise raise NonMonotonicIndex, because every read of the
archive would.
A missing quality column is a refusal, not a default: a value whose
trustworthiness is unknown is not a measurement. assume_quality
is the explicit override, and it is recorded on the archive
(quality_assumed) and printed by every profile of it.
sep, decimal and encoding read a CSV written with another
column separator, decimal mark or text encoding, such as the
;-separated, comma-decimal cp1252 file of a German Excel. Left
None, a CSV reads comma-separated, with a decimal point, as UTF-8.
Rows with different UTC offsets each keep their own. A numeric date
such as 01/02/2026 reads as 1 February day first and 2 January
month first; a column holding one raises SchemaError unless
dayfirst or timestamp_format (a strptime format every row must
match) states the order. A column whose dates read only one way, such
as 3/14/2024 1:05 PM, parses without either.
Raises:
| Type | Description |
|---|---|
SchemaError
|
unreadable input, a CSV header that holds |
NonMonotonicIndex
|
a timestamp precedes the row before it. |
ValueError
|
both |
FileExistsError
|
|
Examples:
A CSV export of the demo flow tag, ingested under a new identity:
>>> import pandas as pd
>>> import tsdive
>>> pd.read_parquet("data/demo/fic101_demo.parquet").to_csv("fic101.csv", index=False)
>>> meta = tsdive.TagMeta(identity=tsdive.TagIdentity("plant1", "FIC101.PV"),
... name="FIC-101 flow", unit_raw="m3/h")
>>> out = tsdive.ingest("fic101.csv", out="archive/plant1/FIC101.PV.parquet", meta=meta)
>>> out.as_posix(), len(pd.read_parquet(out))
('archive/plant1/FIC101.PV.parquet', 562)
An export whose dates read both ways states its order:
>>> pd.DataFrame({"timestamp": ["01/02/2026 08:00:00", "13/02/2026 08:00:00"],
... "value": [61.0, 62.0], "quality": ["GOOD", "GOOD"]}
... ).to_csv("eu.csv", index=False)
>>> tsdive.ingest("eu.csv", out="eu.parquet", meta=meta, tz="Europe/Paris")
Traceback (most recent call last):
...
tsdive.errors.SchemaError: timestamp: '01/02/2026 08:00:00' reads as day 01 of month 02 ...
>>> out = tsdive.ingest("eu.csv", out="eu.parquet", meta=meta, tz="Europe/Paris",
... dayfirst=True)
>>> list(pd.read_parquet(out)["timestamp"].dt.strftime("%Y-%m-%d %H:%M"))
['2026-02-01 07:00', '2026-02-13 07:00']