IECQ006: shared_code_extraction¶
Code is extracted from a completion with inspect_evals' extract_code_block, not a hand-rolled fence pattern.
Category: code_quality ยท Applies to: eval, helper
What it does¶
Flags a string literal holding a triple-backtick fence, or a regex quantifier asking for three backticks, passed as the pattern to re.compile, search, findall, finditer, match, fullmatch or sub, or as a tag to extract_from_tags. As the separator to a split, replace, find, partition or startswith method it counts only with a language label, such as "```python": a bare fence there is as likely a toggle while reading markdown. The literal may be an f-string or a + concatenation, or a name bound to one in the same function or at module level. One diagnostic per function, at its first such call, and one for the module-level calls in a file.
A pattern whose fences are all labelled json is JSON parsing, not code extraction, and is left alone. So is a bare fence, such as a closing fence being stripped, beside one in the same function. A label alternation that names another language, such as (?:json|python)?, still counts. Test files and the module that defines the helper are not read.
The rule applies only where the helper is available: in inspect_evals itself, where a helper package (helper-dirs) defines extract_code_block, and in a repository whose pyproject.toml declares a dependency on inspect-evals. Elsewhere it is skipped.
Why is this bad?¶
Hand-rolled fence patterns disagree on the edge cases: prose before the fence, an unlabelled fence, more than one block, a fence that does not start a line, CRLF line endings. Two evaluations of the same model then extract different code from the same kind of answer, and a bug fixed in one stays in the rest. extract_code_block(completion, language) prefers a block labelled with the requested language over an unlabelled one, returns the first such block, and returns None when there is none, so the caller decides what an answer without a block means. A labelled fence may follow prose on the same line, and one that is never closed runs to the end of the text. extract_code_blocks returns every such block in order, for an evaluation that takes the last block or all of them.
Keep each evaluation's choice of block. accept=defines(name) keeps the first block whenever it defines the function the task asks for, and otherwise takes the first block that does, so it moves results only where the first block could not have passed. Suppress the finding, with a reason that names the rule, only where the evaluation extracts in a way the helpers can't express.
Migrating changes which text is extracted for some completions, so scores can move and a migration needs a comparability bump to the task version and a changelog entry. Compare the old pattern with the helper on the shapes the pattern handled, and keep the old pattern as a fallback for any the helper loses.
Example¶
def find_code(completion: str) -> str:
pattern = re.compile(r"```python\n(.*?)```", re.DOTALL)
matches = pattern.findall(completion)
return matches[0] if matches else completion
Use instead:
from inspect_evals.utils.code import extract_code_block
def find_code(completion: str) -> str:
return extract_code_block(completion, "python") or completion
Where the evaluation takes another block:
from inspect_evals.utils.code import extract_code_blocks
blocks = extract_code_blocks(text, "cpp")
code = next((block for block in reversed(blocks) if "#include" in block), None)
Or, where it extracts in a way the helpers can't express:
blocks = re.findall(r"```(?:cuda|cpp)\n(.*?)```", text, re.DOTALL) # inspect-evals-lint: ignore[IECQ006] -- upstream reads CUDA and C++ blocks in one pass
Suppress on a line with # inspect-evals-lint: ignore[IECQ006] -- <reason> or ignore[shared_code_extraction] -- <reason>; select or ignore it in configuration by either, or by the prefix IECQ.