Skip to content

API reference / Reading and ingest

tsdive.ingest_long

ingest_long(
    source: str | Path,
    *,
    out_dir: str | Path,
    meta_dir: str | Path,
    tag_col: str,
    timestamp_col: str = "timestamp",
    value_col: str = "value",
    quality_col: str = "quality",
    tags: Sequence[str] | None = None,
    tz: str | None = None,
    assume_quality: str | None = None,
    overwrite: bool = False,
    timestamp_format: str | None = None,
    dayfirst: bool = False,
    sep: str | None = None,
    decimal: str | None = None,
    encoding: str | None = None,
) -> list[Path]

Turn a long export, one row per tag and timestamp, into one archive per tag.

The file is read once. tag_col names each row's tag, and the tags are tags when given, else every value of tag_col in order of first appearance. Metadata for a tag comes from meta_dir / f"{safe_filename(tag)}.json" and its archive goes to out_dir / f"{safe_filename(point_id)}.parquet". Each archive holds the tag's rows in file order. Every check runs before the first archive is written, so a refusal leaves out_dir as it was. The file is read as ingest reads it, sep, decimal and encoding included. Timestamps are parsed with one date order for the whole column, and localised tag by tag.

Raises:

Type Description
SchemaError

unreadable input; a missing timestamp, value, tag or quality column or metadata file; a row without a tag; a tag in tags that the column does not hold; naive timestamps without tz; a date that reads day first and month first with no order stated; or an unusable assume_quality value.

NonMonotonicIndex

a tag's timestamp precedes the one before it.

ValueError

two tags or two point ids that share one file name, or both timestamp_format and dayfirst.

FileExistsError

an archive exists and overwrite is False.

Examples:

A long export of the two demo tags, without a quality column:

>>> import pandas as pd
>>> import tsdive
>>> parts = [
...     pd.read_parquet(f"data/demo/{name}_demo.parquet")[["timestamp", "value"]]
...     .assign(tag=tag)
...     for name, tag in (("fic101", "FIC101.PV"), ("tic101", "TIC101.PV"))
... ]
>>> pd.concat(parts).sort_values("timestamp", kind="stable").to_csv(
...     "long.csv", index=False)
>>> _ = tsdive.init_long_meta("long.csv", out_dir="meta", source_id="plant1",
...                           tag_col="tag")
>>> archives = tsdive.ingest_long("long.csv", out_dir="archive/plant1",
...                               meta_dir="meta", tag_col="tag",
...                               assume_quality="GOOD")
>>> [(path.as_posix(), len(pd.read_parquet(path))) for path in archives]
[('archive/plant1/FIC101.PV.parquet', 562), ('archive/plant1/TIC101.PV.parquet', 562)]