Evidence-as-Code: Making AI Governance Verifiable
SGAEIA Research Series — Article 6 Aridio Silva · Independent Researcher, Brazil · ORCID An AI governance claim is only as defensible as the evidence that supports it. A policy may describe intended behavior. A log m
SGAEIA Research Series — Article 6
Aridio Silva · Independent Researcher, Brazil · ORCID
An AI governance claim is only as defensible as the evidence that supports it.
A policy may describe intended behavior. A log may describe an event. An approval record may show that someone clicked a button. None of these, alone, necessarily establishes that an important governance claim is justified.
This is a technical edition of the same public research work published on the SGAEIA homepage and Medium and archived on Zenodo. It reorganizes the discussion for developers and architects without changing the article's thesis, evidence, limitations, or public-disclosure boundary.
Contents
- A governance claim needs a defensible basis
- What Evidence-as-Code means
- Accountability remains important when work is delegated
- Five qualities of useful evidence
- Separate records, observations, and conclusions
- How evidence can mislead
- Respect people and confidential information
- Keep assessment open to revision
- A developer-oriented evidence pattern
- The public contribution and its limits
- Conclusion
- References
- Research and project resources
- License and status
A governance claim needs a defensible basis
An organization may say that its AI system is governed because it has written policies, assigned responsibilities, or introduced an approval process. These are meaningful commitments. Assessing whether a particular claim holds requires attention to what the available evidence actually supports.
Consider a fictional AI assistant helping prepare a purchase. The organization states that important purchases receive appropriate human oversight. A record showing that a purchase occurred does not, by itself, establish that the oversight was meaningful. An approval statement does not establish that the decision was appropriate, that the reviewer had enough context, or that affected people could challenge it.
These are different questions, even when they concern the same event. Assurance is necessarily bounded: a conclusion concerns a particular claim, an identified context, and information with known limitations. A favorable finding about one question should not silently expand into a claim that the entire system is trustworthy.

Figure 1 — Claims and Evidence. A statement about governed behavior and the information supporting its assessment serve different purposes; the illustration does not prescribe a processing sequence. © 2026 Aridio Silva | Project SGAEIA | CC BY 4.0.
What Evidence-as-Code means
Here, Evidence-as-Code refers to an approach in which governance-relevant information is expressed clearly enough to support repeatable assessment, including machine-assisted assessment where appropriate.
The word “code” highlights the possibility of consistent interpretation and review. It does not make software output equivalent to proof, and it does not remove the need for judgment.
NIST OSCAL provides machine-readable representations for security controls and assessment information [1]. Related research examines OSCAL as a candidate format for machine-readable AI-compliance evidence [2]. These are relevant foundations, not evidence that every question about autonomous AI can be reduced to an automated test or that SGAEIA has been validated.
The public SGAEIA contribution here is an interpretive argument: evidence quality should be central to discussions of autonomous-AI governance. This is a research position, not a claim that the term is new, that a universal evidence standard exists, or that a particular implementation has been proven.
MachineReadable ≠ AutomaticallyTrue
Accountability remains important when work is delegated
An AI-assisted activity may involve several people, organizations, or automated services. Distributing the work can make responsibility harder to understand. It does not make accountability less necessary.
For a developer or reviewer, four questions are useful prompts:
- Responsibility: who is answerable for the activity and its consequences?
- Justification: what supports the claim being made about that activity?
- Consequences: what was the observed impact, including adverse effects?
- Uncertainty: what cannot presently be established?
These questions are a non-exhaustive public synthesis, not a validated assessment instrument. They are not a specification of what every system must collect, a model of agent coordination, or a design for authorizing delegated actions. Their value is practical: the presence of several participants must not become an excuse for an unanswerable outcome.

Figure 2 — Questions for Accountable AI. Delegation does not erase accountability; these questions help readers assess an explanation without prescribing internal relationships or mechanisms. © 2026 Aridio Silva | Project SGAEIA | CC BY 4.0.
Five qualities of useful evidence
Evidence is useful in relation to a claim. A large quantity of information can still leave the relevant question unanswered. A carefully bounded body of information may support a modest conclusion without supporting broader assurances.
The article proposes five complementary qualities as an explanatory aid:
- Relevance — the information addresses the question under examination.
- Credibility — there is a defensible basis for relying on the information.
- Context — the circumstances and conditions of the account are understandable.
- Sufficiency — the information supports the stated scope of the conclusion.
- Reviewability — another qualified reviewer can critically examine the basis and limits of the assessment.
These are not a standardized taxonomy, validated scale, lifecycle, or exhaustive test of evidence quality. Their importance varies with the claim and the consequences of being wrong. An assessment can be well documented and still depend on assumptions that deserve challenge.

Figure 3 — Properties of Useful Evidence. The properties describe the quality of an evidential basis; they do not represent software components, architectural layers, or an operational model. © 2026 Aridio Silva | Project SGAEIA | CC BY 4.0.
Separate records, observations, and conclusions
Operational information and governance assessment overlap, but they answer different questions. OpenTelemetry describes traces, metrics, and logs as signals representing system activity [3]. These signals can be useful sources of information. Their presence alone does not establish that a governance claim is justified.
Provenance concerns the origins and history of information. W3C PROV provides a basis for describing provenance and supporting assessments of reliability [4]. Knowing where information came from helps a reviewer reason about it, but does not automatically settle questions about responsibility, appropriateness, or truth.
Keep three ideas distinct:
- Authenticity: whether a record is genuinely attributable to its claimed source.
- Integrity: whether the record has remained unaltered in the relevant respects.
- Truthfulness: whether its account accurately represents what happened.
An authentic, unaltered statement can still be mistaken, incomplete, or misleading. A careful assessment keeps observed information separate from the interpretation placed upon it. Reviewers should be able to distinguish what was reported, what was corroborated, what was inferred, and what remains uncertain.
How evidence can mislead
Some of the most consequential weaknesses arise from the way information is selected or interpreted. A collection of accurate records may still omit a relevant event. A favorable summary may obscure disagreement. An apparently comprehensive account may concern only a narrow set of circumstances.
Important limitations include:
- Missing information: absence may prevent a conclusion, even when the available material appears coherent.
- Selective reporting: presenting only favorable material can produce unjustified confidence.
- Misplaced confidence: authenticity or technical validity may be mistaken for truthfulness.
- Excessive collection: seeking more information can create new privacy and security risks.
These limitations should affect the strength of the conclusion. A gap does not automatically establish misconduct, but neither should it be treated as affirmative evidence that everything was appropriate. An honest assessment may conclude that the available information is insufficient.

Figure 4 — Limits of Assurance. Evidence-based assessment must acknowledge uncertainty and potential harm; the figure is not a lifecycle, threat-control map, or implementation guide. © 2026 Aridio Silva | Project SGAEIA | CC BY 4.0.
Respect people and confidential information
Accountability does not require indiscriminate collection. Recording more information may expose personal data, confidential communications, or intellectual property without improving the answer to the question being assessed.
The objective is a proportionate basis for review. An organization should be able to explain why information is needed, whose interests may be affected by its use, and what limitations apply to the account it provides. Evidence collected for one purpose should not quietly become a justification for unrelated surveillance.
Transparency is compatible with protecting sensitive information. A public explanation can describe the question, the basis of a conclusion, and its limits without revealing every internal detail. Appropriate disclosure requires judgment about both accountability and the consequences of exposure.
MoreEvidence ≠ BetterAssurance
Keep assessment open to revision
AI systems and their contexts change. An earlier assessment may remain historically useful while no longer supporting the same conclusion about a later situation. Confidence should therefore remain connected to the circumstances under which it was earned.
The NIST AI Risk Management Framework provides a voluntary basis for considering trustworthiness and managing AI risks [5]. NIST SP 800-53A provides a complementary foundation for assessing security and privacy controls [6]. These sources support disciplined assessment; neither certifies the SGAEIA research presented here.
Continuous assurance should not be understood as a promise of complete visibility, uninterrupted certainty, or automatic resolution of governance questions. A continuing commitment to assessment includes acknowledging contrary evidence and revisiting earlier conclusions.
A developer-oriented evidence pattern
The following pseudocode is a didactic example, not an SGAEIA implementation, evidence schema, or assurance guarantee:
claim = governance_claim.define(
subject=system_or_activity,
scope=approved_scope,
question="Was the consequential action governed under the approved conditions?"
)
records = evidence_sources.collect_for(claim)
assessment = reviewer.assess(
relevance=records.relevance_to(claim),
credibility=records.source_and_provenance(),
context=records.operating_conditions(),
sufficiency=records.supported_scope(),
reviewability=records.reproducible_basis()
)
result = {
"claim": claim,
"supported": assessment.supported,
"limitations": assessment.uncertainties,
"contrary_evidence": assessment.open_questions,
"reviewed_at": assessment.time
}
evidence_store.record(result)
The important boundary is conceptual rather than syntactic:
- define the claim and its scope;
- collect only information relevant and proportionate to that claim;
- preserve source, context, provenance, and limitations;
- distinguish observation from interpretation;
- record what is supported, what is uncertain, and what remains open;
- allow qualified review and later revision.
Machine assistance can help organize, correlate, and test evidence. It must not silently turn a narrow observation into a universal assurance claim.
The public contribution and its limits
This article offers a vocabulary for discussing why evidence matters and how its quality affects governance claims. The figures present distinctions, questions, properties, and limitations. They do not define how to build an SGAEIA system.
The method is a selected-source narrative synthesis of official guidance on risk management, control assessment, provenance, and observability, together with one closely related research preprint. It is not a systematic literature review: no exhaustive search strategy, study-selection protocol, or quantitative evidence synthesis was conducted.
The questions and properties introduced here are explanatory proposals, not empirically validated measures. The discussion presents no experimental results demonstrating a particular level of security, reliability, or assurance. It makes no claim of certification or legal conformity. An assessment of a concrete deployment would require its own evidence, scope, methods, and qualified judgment.
Conclusion
As AI systems participate in consequential work, organizations need more than confident statements about governance. They need defensible explanations of what the available evidence supports and where it falls short.
Evidence-as-Code places clarity and assessability at the center of that task. Its public value lies in encouraging better questions: Is the information relevant? Is the account credible? Are its conditions clear? Is the conclusion proportionate? Can the basis be critically reviewed?
Accountable AI requires room for scrutiny, disagreement, and revision. Trust is strengthened when claims remain connected to evidence and uncertainty is made visible.
References
- National Institute of Standards and Technology. Open Security Controls Assessment Language (OSCAL).
- Cilla Ugarte, R. et al. Making AI Compliance Evidence Machine-Readable. arXiv:2604.13767v1, 2026. Preprint; DOI.
- OpenTelemetry. Signals.
- Groth, P.; Moreau, L., eds. PROV-Overview: An Overview of the PROV Family of Documents. W3C Working Group Note.
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1, 2023.
- Joint Task Force. Assessing Security and Privacy Controls in Information Systems and Organizations. NIST SP 800-53A Rev. 5.
Research and project resources
- Canonical homepage reading edition
- Article 6 Zenodo DOI
- Original Medium publication
- Academia.edu publication
- SGAEIA homepage
- SGAEIA research artifact
- ORCID — Aridio Silva
- Google Scholar — Aridio Silva
Suggested citation: Silva, Aridio. (2026). Evidence-as-Code: Making AI Governance Verifiable. Zenodo. https://doi.org/10.5281/zenodo.22736488
License and status
The text and original illustrations are licensed under CC BY 4.0, except where otherwise noted. Referenced third-party works retain their own terms.
This DEV Community draft is a technical edition of the same public research work. It is an educational research synthesis, not a systematic review, validated protocol, production certification, legal-compliance determination, or empirical proof of SGAEIA effectiveness. The pseudocode is didactic and does not expose private mechanisms.
© 2026 Aridio Silva | Project SGAEIA | CC BY 4.0
Autonomous AI. Governed by Design. Trusted by Evidence.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.