NAESTRO / EVALUATION

Evidence should travel with the work.

A useful evaluation does more than score an answer. It makes the path inspectable: the admitted context, the actions taken, the controls encountered, the evidence retained and the questions that remain open.

Illustrative comparisonTwo paths. An inspectable difference.
  • Action withheld

    Missing evidence keeps the action from proceeding.

  • Review export admitted

    Added context permits an export for inspection.

A reviewable trajectory

01

What was the request and which context was admitted?

02

Which tools and providers were available at each step?

03

Where did policy allow, pause or refuse the trajectory?

04

What evidence supports the outcome, and what remains unknown?

Compare what changed

Run comparison is most useful when it names the difference. A changed provider, context snapshot, policy decision or tool result can explain why two trajectories diverged. The reviewer should be able to see the changed input and the resulting evidence without treating a polished answer as proof of its own claims.

Missing evidence. Action stops here.
PATH A / REFUSED

Action withheld

The requested publish action is refused because the required evidence is missing. The record preserves the refusal and its reason.

Context checked. Export available for review.
PATH B / REVIEW EXPORT

Evidence admitted

A reviewer supplies the missing context and receives an export for inspection. The example does not imply that the original action was silently published.

A scoped independent-review pilot

Begin with one workflow, a fixed question set and an agreed evidence schema. Separate retrieval quality, report quality, conformance and cost. Record the corpus or fixture version, provider mode, budget and failure accounting before the run; keep the result scoped to that protocol.

The proposed first proof is one agent workflow, one agreed evaluation protocol and one proposed improvement. A second reviewer should be able to reconstruct the recorded failure, compare the baseline with the change, and assess the evidence without rebuilding the investigation.

  1. Agree the boundary. Define the workflow, data location, approved integrations, access permissions and acceptance criteria.
  2. Record the baseline. Retain versions of the model configuration, available context, tools and applicable policies.
  3. Evaluate the change. Preserve proposed modifications, test results, failed checks, refusals and missing evidence alongside successful outcomes.
  4. Hand over the record. Give an authorized reviewer a clear comparison, unresolved questions and the decision about adopting the change.

Integration with an existing runtime and evaluation stack is part of the pilot scope. Data handling and any evidence export must satisfy the agreed environment and access constraints. Reconstructing an execution record does not promise identical regeneration of stochastic model outputs.

Discuss pilot scope, access & licensing

NAESTRO’s benchmark archive is a retained public artifact with its own task and result scope. It is useful evidence for that benchmark, not a general intelligence claim. Open the benchmark archive .

Evidence and trust

Hashing, canonical serialization, issuer authentication and independent witnessing answer different questions. NAESTRO keeps those boundaries visible: a local hash or finite replay is not presented as an independently witnessed certificate.

Propose a scoped evaluation

Bring a real workflow. Start with a bounded question.

Discuss an evaluation