IEBP014: metric_epoch_safety¶
Custom metrics read scores as the epoch reducer leaves them and do not modify them.
Category: best_practices ยท Applies to: eval, helper
What it does¶
Reads every @metric function, including the function it returns and any helper nested in it, and every @score_reducer function. Three things are flagged, one diagnostic per site:
- Narrowing a score value in a metric (error):
.as_int(),.as_bool(), orint(...)of an expression computed from.value, such asint(value["count"])orint(sum(values)).int(len(...))andint()of a comparison are not flagged.isinstance(v, int)orisinstance(v, float)applied directly to a score value is a warning;isinstance(v, (int, float))and the NaN checkisinstance(v, float) and math.isnan(v)are not flagged. - Reading
<x>.score.answer,.explanationor.reasonin a metric (error). - Writing to the scores a metric or reducer was handed (error): an attribute or subscript assignment rooted at its scores parameter, at a loop variable over it, or at a name bound to either. The scores parameter is the first parameter of the function a
@metricor@score_reducerfactory returns, or any nested parameter annotatedSampleScoreorScore. Names are followed in statement order, so rebinding one to a new object, such ass = s.model_copy(deep=True)orcopy.deepcopy(s), ends it. A shallowmodel_copy()orcopy.copy()shares the objects it holds, soc.score.value = 0on one is flagged andc.score = ...is not.
A metric declared @metric(scores="unreduced") receives each epoch's own score, so only the third check applies to it. A plain function that a metric calls is not read.
Why is this bad?¶
Inspect runs the epoch reducer before any metric, even at epochs=1 (_reduced_score in inspect_ai/scorer/_reducer/reducer.py). The score a metric receives is a new one, and its value is whatever the reducer made of the epochs, so a metric cannot assume the shape the scorer wrote. The default mean, and median, pass_at and pass_k, compute a float through value_to_float, so under mean it can be a fraction such as 0.667. at_least gives 1 or 0. mode, majority, max and collect keep the raw values. answer, explanation and reason are kept only when they are equal across all epochs, and are otherwise None. metadata comes from the first epoch.
So int() and .as_int() round a sample that passed two epochs out of three down to 0, .as_bool() rounds it up to True, and a single-type isinstance check is true or false depending on which reducer the run used. A metric that counts score.answer == "win" counts a sample that won one epoch and lost another as neither. All of these pass a single-epoch test and go wrong with --epochs. Writing to the scores is a different hazard: Inspect hands the same score objects to every metric on a scorer, and grouped() hands them to its inner metric more than once, so a later metric sees the changed values.
Example¶
@metric
def win_rate() -> Metric:
def metric(scores: list[SampleScore]) -> float:
return sum(s.score.answer == "win" for s in scores) / len(scores)
return metric
Use instead: give the scorer a value every reducer can average, and read it as a float.
# scorer: Score(value={"win": 1.0 if won else 0.0}, answer=outcome)
@metric
def win_rate() -> Metric:
def metric(scores: list[SampleScore]) -> float:
return sum(s.score.as_dict()["win"] for s in scores) / len(scores)
return metric
See also¶
Suppress on a line with # inspect-evals-lint: ignore[IEBP014] -- <reason> or ignore[metric_epoch_safety] -- <reason>; select or ignore it in configuration by either, or by the prefix IEBP.