Skip to content

API reference / Reading and ingest

tsdive.ingest

ingest(
    source: str | Path,
    *,
    out: str | Path,
    meta: TagMeta,
    timestamp_col: str = "timestamp",
    value_col: str = "value",
    quality_col: str = "quality",
    tz: str | None = None,
    assume_quality: str | None = None,
    overwrite: bool = False,
    timestamp_format: str | None = None,
    dayfirst: bool = False,
    sep: str | None = None,
    decimal: str | None = None,
    encoding: str | None = None,
) -> Path

Turn a CSV or parquet export into a tsdive archive.

Only the three declared columns are carried over; everything else in the export is left behind, so an ingested archive contains exactly what its schema promises. A left-behind column that splits the rows into overlapping series, each in time order, marks an export of several tags. Such an export raises SchemaError when its timestamps repeat or run backwards, or when more than half of its consecutive rows move to another series, as tags on offset clocks do. ingest_long writes one archive per tag from it. The offset-clock rule skips a column of floats, which holds measurements. Timestamps that run backwards otherwise raise NonMonotonicIndex, because every read of the archive would.

A missing quality column is a refusal, not a default: a value whose trustworthiness is unknown is not a measurement. assume_quality is the explicit override, and it is recorded on the archive (quality_assumed) and printed by every profile of it.

sep, decimal and encoding read a CSV written with another column separator, decimal mark or text encoding, such as the ;-separated, comma-decimal cp1252 file of a German Excel. Left None, a CSV reads comma-separated, with a decimal point, as UTF-8.

Rows with different UTC offsets each keep their own. A numeric date such as 01/02/2026 reads as 1 February day first and 2 January month first; a column holding one raises SchemaError unless dayfirst or timestamp_format (a strptime format every row must match) states the order. A column whose dates read only one way, such as 3/14/2024 1:05 PM, parses without either.

Raises:

Type Description
SchemaError

unreadable input, a CSV header that holds ; in one column, bytes that are not encoding, missing column, naive timestamps without tz, a date that reads day first and month first with no order stated, a row that does not match timestamp_format, an unusable assume_quality value, or an export of several tags.

NonMonotonicIndex

a timestamp precedes the row before it.

ValueError

both timestamp_format and dayfirst, an unknown encoding, or a CSV option given for a parquet file.

FileExistsError

out exists and overwrite is False.

Examples:

A CSV export of the demo flow tag, ingested under a new identity:

>>> import pandas as pd
>>> import tsdive
>>> pd.read_parquet("data/demo/fic101_demo.parquet").to_csv("fic101.csv", index=False)
>>> meta = tsdive.TagMeta(identity=tsdive.TagIdentity("plant1", "FIC101.PV"),
...                       name="FIC-101 flow", unit_raw="m3/h")
>>> out = tsdive.ingest("fic101.csv", out="archive/plant1/FIC101.PV.parquet", meta=meta)
>>> out.as_posix(), len(pd.read_parquet(out))
('archive/plant1/FIC101.PV.parquet', 562)

An export whose dates read both ways states its order:

>>> pd.DataFrame({"timestamp": ["01/02/2026 08:00:00", "13/02/2026 08:00:00"],
...               "value": [61.0, 62.0], "quality": ["GOOD", "GOOD"]}
...              ).to_csv("eu.csv", index=False)
>>> tsdive.ingest("eu.csv", out="eu.parquet", meta=meta, tz="Europe/Paris")
Traceback (most recent call last):
...
tsdive.errors.SchemaError: timestamp: '01/02/2026 08:00:00' reads as day 01 of month 02 ...
>>> out = tsdive.ingest("eu.csv", out="eu.parquet", meta=meta, tz="Europe/Paris",
...                     dayfirst=True)
>>> list(pd.read_parquet(out)["timestamp"].dt.strftime("%Y-%m-%d %H:%M"))
['2026-02-01 07:00', '2026-02-13 07:00']