API reference / Reading and ingest
tsdive.ingest_wide
¶
ingest_wide(
source: str | Path,
*,
out_dir: str | Path,
meta_dir: str | Path,
timestamp_col: str = "timestamp",
tags: Sequence[str] | None = None,
quality_suffix: str | None = None,
tz: str | None = None,
assume_quality: str | None = None,
overwrite: bool = False,
timestamp_format: str | None = None,
dayfirst: bool = False,
sep: str | None = None,
decimal: str | None = None,
encoding: str | None = None,
) -> list[Path]
Turn a wide export, one column per tag, into one archive per tag.
The file is read once. Tag columns are tags when given, else every
column that is not timestamp_col and not a quality column; with
quality_suffix a tag's quality column is f"{tag}{quality_suffix}".
Metadata for a tag comes from meta_dir / f"{safe_filename(tag)}.json"
and its archive goes to out_dir / f"{safe_filename(point_id)}.parquet".
Every check runs before the first archive is written, so a refusal
leaves out_dir as it was. The file and its timestamps are read
as ingest reads them, sep, decimal,
encoding, timestamp_format and dayfirst included.
Raises:
| Type | Description |
|---|---|
SchemaError
|
unreadable input; a missing timestamp, tag, quality
or metadata file; naive timestamps without |
NonMonotonicIndex
|
a timestamp precedes the row before it. |
ValueError
|
two tags or two point ids that share one file name, or
both |
FileExistsError
|
an archive exists and |
Examples:
A wide export of the two demo tags, without a quality column:
>>> import pandas as pd
>>> import tsdive
>>> fic = pd.read_parquet("data/demo/fic101_demo.parquet")
>>> tic = pd.read_parquet("data/demo/tic101_demo.parquet")
>>> wide = pd.DataFrame({"ts": fic["timestamp"], "FIC101.PV": fic["value"],
... "TIC101.PV": tic["value"]})
>>> wide.to_csv("export.csv", index=False)
>>> _ = tsdive.init_meta("export.csv", out_dir="meta", source_id="plant1",
... timestamp_col="ts")
>>> archives = tsdive.ingest_wide("export.csv", out_dir="archive/plant1",
... meta_dir="meta", timestamp_col="ts",
... assume_quality="GOOD")
>>> [path.as_posix() for path in archives]
['archive/plant1/FIC101.PV.parquet', 'archive/plant1/TIC101.PV.parquet']