IETS008: custom_metric_tests¶
Every @metric function is tested through the real epoch reducer.
Category: tests ยท Applies to: eval, helper
What it does¶
Finds functions decorated with @metric in the package and checks each is named in a test file that imports the epoch helper module, by default metric_epochs (tests/utils/metric_epochs.py in inspect_evals and the evaluation template). Any import counts: from tests.utils.metric_epochs import ..., from tests.utils import metric_epochs or import tests.utils.metric_epochs. The name must appear as a whole word, so a test for category_win_rate does not cover win_rate. A metric registered under another name with @metric(name=...) is also covered by that name. For an evaluation the search covers tests/<name>/; for a helper package, the whole tests root. One warning per metric. This is a presence check: the test's assertions are the author's.
A hand-written test that runs eval() with epochs= does exercise the reducer, but only for the reducer and values it chose. It is not counted. Moving it to the helper adds the check that equal epochs under mean and mode match the scorer's own values; otherwise suppress with a reason.
Why is this bad?¶
A metric receives the score the epoch reducer leaves, not the one the scorer returned. The default mean turns "C" into 1.0 even at epochs=1, and two passes in three epochs into 0.667. A unit test that hands the metric hand-built scores skips that step, so it passes while the eval reports the wrong number. The helper runs the metric through a real eval: assert_agreeing_epochs_change_nothing checks equal epochs change nothing, and run_metrics takes disagreeing epochs and an expected value.
Example¶
# tests/my_eval/test_my_eval.py: hand-built scores only
def test_win_rate():
assert win_rate()([SampleScore(score=Score(value="C"))]) == 1.0
Use instead:
from tests.utils.metric_epochs import assert_agreeing_epochs_change_nothing, run_metrics
def test_win_rate_through_the_reducer():
assert_agreeing_epochs_change_nothing([win_rate()], [CORRECT, INCORRECT])
assert run_metrics([win_rate()], [[CORRECT, CORRECT, INCORRECT]])["win_rate"] == pytest.approx(2 / 3)
Options¶
custom_metric_tests.helper-module: the module a test must import. A dotted path is matched by its last part. Defaultmetric_epochs.
See also¶
Suppress on a line with # inspect-evals-lint: ignore[IETS008] -- <reason> or ignore[custom_metric_tests] -- <reason>; select or ignore it in configuration by either, or by the prefix IETS.