Skip to content

API reference / Reading and ingest

tsdive.ingest_wide

ingest_wide(
    source: str | Path,
    *,
    out_dir: str | Path,
    meta_dir: str | Path,
    timestamp_col: str = "timestamp",
    tags: Sequence[str] | None = None,
    quality_suffix: str | None = None,
    tz: str | None = None,
    assume_quality: str | None = None,
    overwrite: bool = False,
    timestamp_format: str | None = None,
    dayfirst: bool = False,
    sep: str | None = None,
    decimal: str | None = None,
    encoding: str | None = None,
) -> list[Path]

Turn a wide export, one column per tag, into one archive per tag.

The file is read once. Tag columns are tags when given, else every column that is not timestamp_col and not a quality column; with quality_suffix a tag's quality column is f"{tag}{quality_suffix}". Metadata for a tag comes from meta_dir / f"{safe_filename(tag)}.json" and its archive goes to out_dir / f"{safe_filename(point_id)}.parquet". Every check runs before the first archive is written, so a refusal leaves out_dir as it was. The file and its timestamps are read as ingest reads them, sep, decimal, encoding, timestamp_format and dayfirst included.

Raises:

Type Description
SchemaError

unreadable input; a missing timestamp, tag, quality or metadata file; naive timestamps without tz; a date that reads day first and month first with no order stated; or quality_suffix given together with assume_quality.

NonMonotonicIndex

a timestamp precedes the row before it.

ValueError

two tags or two point ids that share one file name, or both timestamp_format and dayfirst.

FileExistsError

an archive exists and overwrite is False.

Examples:

A wide export of the two demo tags, without a quality column:

>>> import pandas as pd
>>> import tsdive
>>> fic = pd.read_parquet("data/demo/fic101_demo.parquet")
>>> tic = pd.read_parquet("data/demo/tic101_demo.parquet")
>>> wide = pd.DataFrame({"ts": fic["timestamp"], "FIC101.PV": fic["value"],
...                      "TIC101.PV": tic["value"]})
>>> wide.to_csv("export.csv", index=False)
>>> _ = tsdive.init_meta("export.csv", out_dir="meta", source_id="plant1",
...                      timestamp_col="ts")
>>> archives = tsdive.ingest_wide("export.csv", out_dir="archive/plant1",
...                               meta_dir="meta", timestamp_col="ts",
...                               assume_quality="GOOD")
>>> [path.as_posix() for path in archives]
['archive/plant1/FIC101.PV.parquet', 'archive/plant1/TIC101.PV.parquet']