Action withheld
The requested publish action is refused because the required evidence is missing. The record preserves the refusal and its reason.
NAESTRO / EVALUATION
A useful evaluation does more than score an answer. It makes the path inspectable: the admitted context, the actions taken, the controls encountered, the evidence retained and the questions that remain open.
Missing evidence keeps the action from proceeding.
Added context permits an export for inspection.
What was the request and which context was admitted?
Which tools and providers were available at each step?
Where did policy allow, pause or refuse the trajectory?
What evidence supports the outcome, and what remains unknown?
Run comparison is most useful when it names the difference. A changed provider, context snapshot, policy decision or tool result can explain why two trajectories diverged. The reviewer should be able to see the changed input and the resulting evidence without treating a polished answer as proof of its own claims.
The requested publish action is refused because the required evidence is missing. The record preserves the refusal and its reason.
A reviewer supplies the missing context and receives an export for inspection. The example does not imply that the original action was silently published.
Begin with one workflow, a fixed question set and an agreed evidence schema. Separate retrieval quality, report quality, conformance and cost. Record the corpus or fixture version, provider mode, budget and failure accounting before the run; keep the result scoped to that protocol.
The proposed first proof is one agent workflow, one agreed evaluation protocol and one proposed improvement. A second reviewer should be able to reconstruct the recorded failure, compare the baseline with the change, and assess the evidence without rebuilding the investigation.
Integration with an existing runtime and evaluation stack is part of the pilot scope. Data handling and any evidence export must satisfy the agreed environment and access constraints. Reconstructing an execution record does not promise identical regeneration of stochastic model outputs.
Discuss pilot scope, access & licensingNAESTRO’s benchmark archive is a retained public artifact with its own task and result scope. It is useful evidence for that benchmark, not a general intelligence claim. Open the benchmark archive .
Hashing, canonical serialization, issuer authentication and independent witnessing answer different questions. NAESTRO keeps those boundaries visible: a local hash or finite replay is not presented as an independently witnessed certificate.
Propose a scoped evaluationBring a real workflow. Start with a bounded question.
Discuss an evaluation