tsdive.eval.group_holdout
¶
group_holdout(
items: DataFrame,
*,
group_col: str,
stratum_col: str,
n_folds: int = 5,
seed: int = 42,
) -> GroupSplit
Assign whole groups to n_folds folds, balancing strata.
Groups are taken largest first - ties broken by a seeded SHA-256 rank
of the group key, so the order is identical on any platform - and each
is placed in the fold that leaves the per-stratum counts least spread.
A group is never divided: a well contributing 500 instances lands in
exactly one fold, and the fold it lands in is 500 instances larger
than its siblings. That imbalance is the shape of the data, and
GroupSplit.fold_sizes reports it rather than
hiding it behind a row-level cut.
Examples:
>>> import pandas as pd
>>> from tsdive.eval import group_holdout
>>> items = pd.DataFrame({"asset": ["A", "A", "B", "C", "C", "C"],
... "fault": ["x", "y", "x", "y", "x", "x"]})
>>> split = group_holdout(items, group_col="asset", stratum_col="fault",
... n_folds=2, seed=0)
>>> split.groups(0), split.groups(1), split.fold_sizes
(('C',), ('A', 'B'), (3, 3))
>>> split.test_indices(items, 0), split.train_indices(items, 0)
([3, 4, 5], [0, 1, 2])
>>> sorted(split.strata_present(0))
['x', 'y']