API reference / Reading and ingest
tsdive.ingest_long
¶
ingest_long(
source: str | Path,
*,
out_dir: str | Path,
meta_dir: str | Path,
tag_col: str,
timestamp_col: str = "timestamp",
value_col: str = "value",
quality_col: str = "quality",
tags: Sequence[str] | None = None,
tz: str | None = None,
assume_quality: str | None = None,
overwrite: bool = False,
timestamp_format: str | None = None,
dayfirst: bool = False,
sep: str | None = None,
decimal: str | None = None,
encoding: str | None = None,
) -> list[Path]
Turn a long export, one row per tag and timestamp, into one archive per tag.
The file is read once. tag_col names each row's tag, and the tags
are tags when given, else every value of tag_col in order of
first appearance. Metadata for a tag comes from
meta_dir / f"{safe_filename(tag)}.json" and its archive goes to
out_dir / f"{safe_filename(point_id)}.parquet". Each archive holds
the tag's rows in file order. Every check runs before the first
archive is written, so a refusal leaves out_dir as it was.
The file is read as ingest reads it, sep,
decimal and encoding included. Timestamps are parsed with one
date order for the whole column, and localised tag by tag.
Raises:
| Type | Description |
|---|---|
SchemaError
|
unreadable input; a missing timestamp, value, tag or
quality column or metadata file; a row without a tag; a tag in
|
NonMonotonicIndex
|
a tag's timestamp precedes the one before it. |
ValueError
|
two tags or two point ids that share one file name, or
both |
FileExistsError
|
an archive exists and |
Examples:
A long export of the two demo tags, without a quality column:
>>> import pandas as pd
>>> import tsdive
>>> parts = [
... pd.read_parquet(f"data/demo/{name}_demo.parquet")[["timestamp", "value"]]
... .assign(tag=tag)
... for name, tag in (("fic101", "FIC101.PV"), ("tic101", "TIC101.PV"))
... ]
>>> pd.concat(parts).sort_values("timestamp", kind="stable").to_csv(
... "long.csv", index=False)
>>> _ = tsdive.init_long_meta("long.csv", out_dir="meta", source_id="plant1",
... tag_col="tag")
>>> archives = tsdive.ingest_long("long.csv", out_dir="archive/plant1",
... meta_dir="meta", tag_col="tag",
... assume_quality="GOOD")
>>> [(path.as_posix(), len(pd.read_parquet(path))) for path in archives]
[('archive/plant1/FIC101.PV.parquet', 562), ('archive/plant1/TIC101.PV.parquet', 562)]