Published evaluation protocol
How a claim earns its place here.
Each future automation result should be tied to a frozen input set, expected outcomes, a defined scorer, recorded exceptions, and a named human approval point. Failed and unsupported cases belong in the record.
01Define the task boundary
State the exact input, output, intended user, supported file or data versions, prohibited uses, and consequence of a missed issue.
02Freeze representative cases
Version the test set and include ordinary examples, shape-family variation, incomplete inputs, ambiguous conditions, and known edge cases without exposing confidential project data.
03Establish expected results
Create the reference result independently of the automation where practical. Record the source, assumptions, units, reviewer, and any cases without a single defensible answer.
04Run and preserve outputs
Record the software and ruleset version, runtime environment, input hash, warnings, raw output, elapsed time, and any manual intervention.
05Score every exception
Classify correct detections, incorrect flags, misses, unsupported cases, and reviewer overrides. Do not remove difficult cases merely because they lower a headline result.
06Publish context and limits
Report sample size, method, metric definitions, results, failures, version, date, reviewer role, conflicts of interest, and whether the result was internal or independently reproduced.