Skip to content

IEBP011: shuffle_seeded

A dataset loader that shuffles samples is given a seed.

Category: best_practices ยท Applies to: eval, helper

What it does

Warns on each shuffle= argument that resolves to true when the call passes no seed=, or a seed= that resolves to None. Values are resolved as shuffle_choices_seeded resolves them, and a value the rule can't resolve is taken at face value. A call with **kwargs and no seed= is taken to pass one.

Any call passing shuffle= by keyword is read, and shuffle and seed may be passed positionally to Inspect's hf_dataset, csv_dataset, json_dataset and file_dataset and inspect_evals' load_csv_dataset and load_json_dataset.

A dataset's .shuffle() method is read too, and warned on when it passes no seed or one that resolves to None. A receiver named random shuffles a list, not a dataset, and is not read. As in shuffle_choices_seeded, a call inside an if that the defaults skip, such as if shuffle: with shuffle=False, is not flagged.

inspect_evals' shuffle_and_seed(shuffle) reads shuffle as inspect reads shuffle_choices: an int is the seed, True is unseeded and False is off. A call to it is flagged when its argument resolves to True, so shuffle: bool | int = 42 passes and = True does not. The loader it feeds is not flagged again.

A dataset helper that takes shuffle= but seeds internally is flagged where it is called without a seed; suppress that finding with the reason. When the value comes from a parameter default, the diagnostic points at the default.

The rule asks for a seed, not for the shuffle to go. Shuffling by default is often right: a dataset whose records come grouped by category, subject or source needs it, or a --limit run evaluates only the first groups. Whether the records are grouped depends on the data, which the rule does not read. Seeding changes which samples a --limit run selects. Removing a shuffle task parameter breaks -T shuffle=... invocations, so it needs a task version bump.

This is a warning rather than an error: an unseeded shuffle changes only the order of the samples, not the samples themselves.

Why is this bad?

Sample order decides which samples --limit picks. Without a seed, two runs with the same limit evaluate different samples, and neither log records the order used, so their scores can't be compared. A seed keeps the benefit of the shuffle, a --limit run that spans the whole dataset, and makes it the same subset each time.

Where the records are in no meaningful order, the shuffle can go instead. inspect eval --sample-shuffle <seed> shuffles at run time and records the seed in the log's eval config.

Example

@task
def my_eval(shuffle: bool = True):
    return Task(dataset=hf_dataset("org/data", split="test", shuffle=shuffle))

Use instead:

@task
def my_eval(shuffle: bool | int = 42):
    should_shuffle, seed = shuffle_and_seed(shuffle)
    return Task(dataset=hf_dataset("org/data", split="test", shuffle=should_shuffle, seed=seed))

See also

Suppress on a line with # inspect-evals-lint: ignore[IEBP011] -- <reason> or ignore[shuffle_seeded] -- <reason>; select or ignore it in configuration by either, or by the prefix IEBP.