Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 10 min read

From Local RAG to a Verifiable Knowledge Proof on Ethereum Sepolia

From Local RAG to a Verifiable Knowledge Proof on Ethereum Sepolia How I connected Docker, Ollama, Qdrant, GitHub, a MyZubster Knowledge Card and an on-chain SHA-256 proof Building an AI project is one thing. Being able

From Local RAG to a Verifiable Knowledge Proof on Ethereum Sepolia
How I connected Docker, Ollama, Qdrant, GitHub, a MyZubster Knowledge Card and an on-chain SHA-256 proof
Building an AI project is one thing.
Being able to show what was actually built, what was tested, where the evidence is, and whether a documented artifact has changed is another.
During the development of my myzubster-mvp project, I wanted to test a simple idea:
Can we connect a person's documented technical activity to GitHub evidence, an AI/RAG system, a public Knowledge Card, and finally a cryptographic proof recorded on a blockchain?

This experiment became a small end-to-end proof-of-concept involving:

  • Docker and Docker Compose
  • Ollama
  • Qdrant
  • Open WebUI
  • a local RAG pipeline
  • GitHub commits
  • MyZubster Knowledge Cards
  • SHA-256
  • a Solidity smart contract
  • Ethereum Sepolia
  • a Knowledge Graph The goal was not to put personal knowledge itself on a blockchain. The goal was to build a verifiable chain between a documented knowledge artifact and a cryptographic commitment to that artifact.
  • Starting point: my local AI/RAG environment The experiment started inside my public project: Repository: https://github.com/nicolaususnicola-lgtm/myzubster-mvp The local stack combines MyZubster's API with Docker, Ollama, Qdrant and Open WebUI. At a simplified level: Documents / observations ↓ Ingestion ↓ Embeddings ↓ Qdrant ↓ Relevant context ↓ Local LLM through Ollama ↓ MyZubster AI response

One of the important principles behind the project is evidence first.
The AI should not simply generate plausible answers.
Where authoritative structured information exists, that information should remain the source of truth. RAG and the local language model can then operate over the evidence retrieved from the system.

  1. Improving the ingestion pipeline During testing I found that document chunking could be improved. Instead of blindly cutting text, I modified the ingestion process so that it tries to respect paragraph and line boundaries while maintaining overlap between chunks. The change is documented here: Chunking commit: https://github.com/nicolaususnicola-lgtm/myzubster-mvp/commit/ea80799 After the modification, I performed a new ingestion. The local test resulted in: 46 knowledge chunks loaded into Qdrant 1 empty document skipped

I also tested retrieval through the MyZubster AI endpoint:
/api/ai/ask

During this work I increased the amount of retrieved AI context from 1 to 5 so that the expected GitHub/project information could be recovered more reliably.
This gave me the first part of the evidence chain:
Technical activity
↓
GitHub commit
↓
Local ingestion
↓
Qdrant
↓
RAG retrieval

  1. Turning the work into a Knowledge Card The next step was to describe the work in a public MyZubster Knowledge Card. The card is: β€œProve Docker e chat AI del progetto myzubster-mvp” Public Knowledge Card: https://www.myzubster.com/knowledge-card?id=6abaaefb3a7460c4574a45fd The card connects the activity to evidence such as the repository, Docker tests, the chunking modification and the later blockchain experiments. This creates another layer: N4K48 ↓ Technical activity ↓ GitHub evidence ↓ Knowledge Card

But at this point there was still an interesting question.
How could I create a cryptographic reference to the Knowledge Card?

  1. Proof v1: anchoring the Knowledge Card URL The first experiment was intentionally simple. I calculated the SHA-256 of the Knowledge Card URL string. That digest was passed to a small Solidity contract deployed on Ethereum Sepolia. The contract was: // SPDX-License-Identifier: MIT pragma solidity ^0.8.20;

contract MyZubsterProof {
bytes32 public knowledgeHash;
address public creator;
uint256 public timestamp;

constructor(bytes32 _knowledgeHash) {
    knowledgeHash = _knowledgeHash;
    creator = msg.sender;
    timestamp = block.timestamp;
}

}

The contract stores three values:
knowledgeHash
creator
timestamp

Proof v1 contract
https://sepolia.etherscan.io/address/0xabCF68e97a32eCa503942A563FF16F209ed45d11#code
Deployment transaction
https://sepolia.etherscan.io/tx/0x09dddd29aca76c9a425ba9cb45cefb1bfbe203b9628f5f9526e4a281012d4f20
GitHub evidence
https://github.com/nicolaususnicola-lgtm/myzubster-mvp/commit/2df5b39
The recorded value was:
0x15b21c4f189259f143f6c946ac866001d88f5bc7aa371cac7cbcc1ce66b43685

The contract source for this first proof was publicly verified on Etherscan as an Exact Match.
But there was an important limitation.
Proof v1 did not hash the Knowledge Card content
It hashed the URL string.
That distinction matters.
The first experiment therefore demonstrated that I could associate a blockchain record with the identifier/location of the Knowledge Card.
It did not cryptographically commit to the complete contents of the card.
So we moved to a second experiment.

  1. Proof v2: anchoring the content For Proof v2, the objective changed. Instead of: Knowledge Card URL ↓ SHA-256

I wanted:
Knowledge Card
↓
Canonical payload
↓
SHA-256
↓
Ethereum Sepolia

The content used for the proof was stored in the GitHub repository so that the artifact being referenced is publicly inspectable.
Canonical payload:
https://github.com/nicolaususnicola-lgtm/myzubster-mvp/blob/main/proofs/knowledge-card-6abaaefb3a7460c4574a45fd-v1.json
The corresponding GitHub commit is:
https://github.com/nicolaususnicola-lgtm/myzubster-mvp/commit/ecefd81c5c9da0be15c99aeeb83878480abd60a8

  1. Calculating the SHA-256 The SHA-256 used for Proof v2 is: 6097e05866bafceec24663d2638cb1dae5742ac78284abbfd45cc9c3b0bfb845

As a Solidity bytes32 value:
0x6097e05866bafceec24663d2638cb1dae5742ac78284abbfd45cc9c3b0bfb845

The important point is that the proof refers to the exact payload bytes committed to the repository.
Anyone with those bytes can independently calculate their SHA-256 and compare the result with the value used in the proof.
For example:
sha256sum proofs/knowledge-card-6abaaefb3a7460c4574a45fd-v1.json

Or with Python:
import hashlibfrom pathlib import Pathdata = Path( "proofs/knowledge-card-6abaaefb3a7460c4574a45fd-v1.json").read_bytes()print(hashlib.sha256(data).hexdigest())

The expected result is:
6097e05866bafceec24663d2638cb1dae5742ac78284abbfd45cc9c3b0bfb845

  1. Registering Proof v2 on Ethereum Sepolia A new MyZubsterProof instance was deployed for Proof v2. Contract https://sepolia.etherscan.io/address/0x21787249Df054132093FcF09bB914C0CCC539390 Deployment transaction https://sepolia.etherscan.io/tx/0xc837ba3f3046f3712e3eba4b81107cb22939b61300f2cab3cfe5bbb7b3319ded Network: Ethereum Sepolia Chain ID: 11155111

After deployment I reconnected to the contract and called:
knowledgeHash()

The returned value was:
0x6097e05866bafceec24663d2638cb1dae5742ac78284abbfd45cc9c3b0bfb845

This matches the digest documented for the committed payload.
That gives us the core verification relationship:
GitHub payload
β”‚
β”‚ SHA-256
β–Ό
6097e05866bafce...
β”‚
β”‚
β–Ό
MyZubsterProof
Ethereum Sepolia
β”‚
β”‚ knowledgeHash()
β–Ό
6097e05866bafce...

  1. Making the proof reproducible A blockchain transaction by itself is not enough for a useful technical experiment. Someone examining the project also needs to understand:
  2. what artifact was hashed;
  3. which digest was expected;
  4. where the artifact is stored;
  5. which contract contains the commitment;
  6. which transaction deployed it;
  7. how to reproduce the hash calculation. For this reason I added dedicated Proof v2 documentation to GitHub: https://github.com/nicolaususnicola-lgtm/myzubster-mvp/blob/main/proofs/SEPOLIA_PROOF_V2.md Documentation commit: https://github.com/nicolaususnicola-lgtm/myzubster-mvp/commit/405467e8d7b4e55ea9f18044a00aa66804f3ada3 This turns the blockchain experiment into something closer to a reproducible evidence record rather than an isolated transaction hash.
  8. Updating the main README One problem remained. The evidence existed, but somebody arriving on the repository homepage might never discover it. So I added a dedicated section to the main README: Knowledge Card β†’ Ethereum Sepolia Proof v2 The README now exposes the verification path directly: N4K48 ↓ Knowledge Card ↓ canonical payload ↓ SHA-256 ↓ Proof v2 Sepolia ↓ GitHub documentation

Repository:
https://github.com/nicolaususnicola-lgtm/myzubster-mvp
README update commit:
https://github.com/nicolaususnicola-lgtm/myzubster-mvp/commit/a4b54b86fe1066684c8604a83b3e32656b7429d7
This is important from a usability perspective.
Evidence that nobody can find is much less useful than evidence that can be navigated.

  1. The complete evidence chain At this stage the experiment can be represented as: N4K48 β”‚ β–Ό Technical activity β”‚ β–Ό GitHub commits β”‚ β–Ό Local AI / RAG tests Docker + Ollama Qdrant + API β”‚ β–Ό Knowledge Card β”‚ β–Ό Canonical payload β”‚ SHA-256 β”‚ β–Ό MyZubsterProof v2 Ethereum Sepolia β”‚ β–Ό GitHub documentation β”‚ β–Ό Knowledge Graph

This is the part of the experiment I find most interesting.
The blockchain is only one component of the evidence chain.
GitHub provides the development history and artifact.
The Knowledge Card provides the human-readable description.
The canonical payload provides the artifact to hash.
SHA-256 provides the cryptographic fingerprint.
Ethereum Sepolia provides a public location for the commitment.
The Knowledge Graph can then provide navigation between those elements.

  1. What did we actually prove? This distinction is essential. The experiment provides evidence that a particular payload corresponds to a particular SHA-256 digest and that this digest was used in the Sepolia proof. It therefore creates a cryptographic link between: documented artifact ↕ cryptographic digest ↕ on-chain record

If the bytes of the artifact are changed, recalculating SHA-256 will normally produce a different digest, so it will no longer match the recorded commitment.
But this does not automatically prove that every statement inside the Knowledge Card is true.
For example, putting the hash of a statement such as:
"I am an expert in X"

on a blockchain does not prove expertise in X.
It proves something much narrower and more technically useful:
this particular data can be associated with this cryptographic fingerprint and compared against the recorded commitment.

That is why MyZubster keeps a distinction between:
CLAIM
EVIDENCE
INTEGRITY
VERIFICATION

They are related concepts, but they are not the same thing.

  1. Why not store the entire Knowledge Card on-chain? Because that is not necessary for this experiment. Instead of storing the whole document on Ethereum, we can store a fixed-size digest: Knowledge Card content ↓ SHA-256 ↓ 32 bytes ↓ Blockchain

The larger artifact can remain in an appropriate external repository or storage layer.
The blockchain only needs the cryptographic commitment.
This approach is simpler and avoids treating a blockchain as general-purpose document storage.

  1. From a Knowledge Card to a Knowledge Graph The next stage is to make the relationship navigable. The intended graph is: N4K48 β”‚ β”œβ”€β”€ Knowledge Card β”‚ β”‚ β”‚ β”œβ”€β”€ GitHub evidence β”‚ β”‚ β”‚ └── Canonical payload β”‚ β”‚ β”‚ β–Ό β”‚ SHA-256 β”‚ β”‚ β”‚ β–Ό β”‚ Proof v2 Sepolia β”‚ β”‚ β”‚ β”œβ”€β”€ Contract β”‚ └── Transaction β”‚ └── Project: myzubster-mvp

The idea is that a visitor should not simply see the statement:
N4K48 worked on a local RAG system.

They should be able to navigate from that statement toward the underlying evidence.
This is the direction being tested with the MyZubster Knowledge Graph.

  1. AI is useful here, but it is not the authority Another lesson from this experiment concerns AI. The local model is useful for:
  2. interpreting questions;
  3. summarizing retrieved context;
  4. navigating documentation;
  5. connecting relevant evidence;
  6. presenting information to a user. But the language model should not become the authoritative source for facts already represented structurally. For example: GitHub commit β†’ evidence JSON payload β†’ artifact SHA-256 β†’ fingerprint Ethereum contract β†’ commitment Qdrant β†’ retrieval LLM β†’ interpretation

The AI operates around the evidence, not instead of it.
This is one of the central ideas behind the evidence-first architecture I am exploring in MyZubster.

  1. What I learned The most important lesson was that verifiability is not created simply by β€œusing blockchain.” A useful evidence system requires several layers to work together: Human-readable description + source evidence + stable artifact + cryptographic digest + public commitment + reproducible verification + clear navigation

Remove the documentation and the hash becomes difficult to interpret.
Remove the original artifact and the digest becomes difficult to reproduce.
Remove the distinction between claims and proofs and it becomes easy to overstate what the technology demonstrates.
The interesting part is therefore not any single technology.
It is the connection between them.

  1. Public references Everything used for this experiment can be inspected through the following starting points: MyZubster MVP repository https://github.com/nicolaususnicola-lgtm/myzubster-mvp Knowledge Card https://www.myzubster.com/knowledge-card?id=6abaaefb3a7460c4574a45fd Chunking improvement https://github.com/nicolaususnicola-lgtm/myzubster-mvp/commit/ea80799 Proof v1 GitHub commit https://github.com/nicolaususnicola-lgtm/myzubster-mvp/commit/2df5b39 Proof v1 Sepolia contract https://sepolia.etherscan.io/address/0xabCF68e97a32eCa503942A563FF16F209ed45d11#code Canonical payload for Proof v2 https://github.com/nicolaususnicola-lgtm/myzubster-mvp/blob/main/proofs/knowledge-card-6abaaefb3a7460c4574a45fd-v1.json Canonical payload commit https://github.com/nicolaususnicola-lgtm/myzubster-mvp/commit/ecefd81c5c9da0be15c99aeeb83878480abd60a8 Proof v2 documentation https://github.com/nicolaususnicola-lgtm/myzubster-mvp/blob/main/proofs/SEPOLIA_PROOF_V2.md Proof v2 documentation commit https://github.com/nicolaususnicola-lgtm/myzubster-mvp/commit/405467e8d7b4e55ea9f18044a00aa66804f3ada3 Proof v2 Ethereum Sepolia contract https://sepolia.etherscan.io/address/0x21787249Df054132093FcF09bB914C0CCC539390 Proof v2 deployment transaction https://sepolia.etherscan.io/tx/0xc837ba3f3046f3712e3eba4b81107cb22939b61300f2cab3cfe5bbb7b3319ded README integration commit https://github.com/nicolaususnicola-lgtm/myzubster-mvp/commit/a4b54b86fe1066684c8604a83b3e32656b7429d7 Conclusion This started as a local Docker and AI/RAG test. It gradually became something broader: local development ↓ documented evidence ↓ Knowledge Card ↓ canonical artifact ↓ cryptographic fingerprint ↓ public blockchain commitment ↓ navigable Knowledge Graph

The experiment does not turn a claim into truth simply because a hash exists on Ethereum.
What it does demonstrate is a practical architecture for making digital knowledge records more traceable, reproducible and tamper-evident.
And for me, that is the interesting direction:
not replacing trust with blockchain or AI, but making the evidence behind a digital claim easier to inspect.
N4K48 / Nicola
Building and testing myzubster-mvp

ai #rag #ethereum #opensource

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.