Dev.to AI 🤖 Ai 👁 0 📖 6 min read

How I built my own AI tools for a small NGO instead of buying software

Small nonprofits have the same paperwork as companies and none of the budget. When I started running one, the problems were familiar: hundreds of emails a month, documents arriving as PDFs, Word files and phone scans, de

How I built my own AI tools for a small NGO instead of buying software

 Small nonprofits have the same paperwork as companies and none of the budget. When I started running one, the problems were familiar: hundreds of emails a month, documents arriving as PDFs, Word files and phone scans, deadlines spread across inboxes, and a small team trying to remember which reply belonged to which file.

The software on the market was either expensive, generic, or built for sales teams. I've spent about twenty years in project and delivery management, so I did what I'd do on any project: broke the work into pieces and built the pieces myself. This time an AI assistant does most of the work, and I supervise.

This post covers the architecture, the parts that failed, and the rules I added after they failed.

Start from the work, not from the model

"We have an AI subscription, what can we do with it?" is the wrong first question. I started by writing down every repetitive task in a normal week and roughly how long each one took. Four came out on top:

  1. Reading incoming mail and working out which file it belongs to.
  2. Turning documents into something searchable.
  3. Tracking deadlines and replies so nothing gets lost.
  4. Finding "that document from three months ago".

None of these needs intelligence in the impressive sense. They need consistency, memory and patience, which an assistant can provide if you give it structure and tools.

The architecture

The setup is an AI assistant (I use Claude) connected to a set of small tools through the Model Context Protocol. MCP is an open standard: each tool runs as a small server that exposes a few functions, and the assistant decides when to call them. A tool does one job. The assistant combines them.

These are the servers I run.

Mail and ticket connector

This one reads the shared inbox, searches threads, downloads attachments and works with the project tracker (Jira). Each incoming message gets linked to the right ticket, with its attachments and a comment that summarizes what changed.

It can also create drafts and send email, and that's exactly why sending is locked behind human approval. The assistant drafts. A person reads and approves. More on this below, because it was the most important design decision.

The same server holds a small deadline register: add, update, supersede, remove and check. Every outgoing request gets a due date calculated from the rules that apply to it. When a date passes, the file comes back for a decision.

Document converter

Everything that arrives (PDF, DOCX, XLSX, PPTX, HTML, ZIP) goes through a converter to Markdown before the model touches it. Plain text is cheaper to process, easier to search and easier to quote exactly.

The exception is images. Modern models read a photo of a document better than OCR does, so images go straight to the model.

Private knowledge base

All files, tickets, documents and replies are indexed in a vector database running on my own machine, with metadata such as source, ticket key and date. The assistant has find, store, update and reindex operations, plus search by entity.

The rule is simple: for any question about our own work, the assistant searches the knowledge base first, and only after that the tracker or the web. Answers have to cite the chunk they came from. This one rule removed most hallucinated answers about our own history, because the model stops guessing when the real document is one call away.

Browser automation

The assistant is told never to try to log in on the anonymous profile. Single-page apps get an explicit wait after each search, because an empty result two seconds after submit is a false negative, not a sign the data doesn't exist. Before long runs, a health check navigates to the authenticated sources and confirms the session is still alive. Having cookies doesn't prove you're logged in.

Skills: writing the process down

The tools were maybe 40% of the value. The rest came from writing down how we work as reusable instructions the assistant can load, which Claude calls "skills". Each one is a Markdown file with a trigger and a procedure:

  • one processes the morning's mail end to end: link, compare the reply with the original request point by point, update the deadline, recommend the next step;
  • one audits an outgoing letter before it's sent;
  • one runs a decision through five opposing perspectives (contrarian, first principles, expansionist, outsider, executor) plus a peer review before we commit;
  • one looks for what's wrong with the previous answer and fixes it.

Writing them forced us to make our process explicit, and that was worth doing for its own sake. A good share of the "AI problems" I fixed turned out to be "we never agreed how to do this" problems.

Verification

AI assistants are confident even when they're wrong. Early on, mine:

  • reported "comment added" when the tool call had failed silently;
  • cited rules from memory instead of the version in force;
  • treated an empty search result as proof that something didn't exist;
  • filled gaps with plausible details.

So verification became part of the design instead of an afterthought.

Every claimed action needs the tool's response. "Email sent" has to come with the message ID the API returned. If there's no tool output, the action didn't happen.

References are checked against the current source. Anything that cites a rule is checked against the official, up-to-date text, not the model's memory. The verified texts are stored once and reused, so the same source isn't fetched again for every document.

Negative results need a control. If a search returns nothing, the same search must find a case we know exists. Otherwise "not found" is a property of the method, not a fact.

There's a verification loop with a cap. After important work, a verification pass looks for concrete problems (unconfirmed actions, unsourced claims, contradictions, unstated assumptions), fixes them in place and runs again until nothing changes. The cap is 5 passes for tool-backed actions and 2 for pure text judgment, so it can't loop forever.

Humans approve external effects. Sending, publishing and filing need explicit approval. The assistant can prepare everything up to the last click.

Building this roughly doubled the effort for each workflow. It's also the only reason I trust the output.

What changed

The measurable gain is time. Sorting a morning's mail, matching it to files and updating deadlines used to take a morning. Now it's a few minutes of review.

The less obvious gain is memory. Nothing depends on one person remembering where a document is or when something is due.

The costs are real too. The assistant still needs supervision. Tools break when a website or an API changes. Some days doing the task by hand is faster than explaining it.

If you want to build something similar

  1. List your repetitive tasks and time them. Automate the top two, not everything.
  2. Keep your data where you control it. A local vector database is cheap and solves most privacy questions, which matters if you handle personal data under GDPR.
  3. Build small tools that do one thing, and let the assistant combine them.
  4. Write the process down as instructions. That's where most of the value is.
  5. Don't skip verification or human approval for anything that leaves the building.

I now help small businesses and independent professionals build the same kind of setup for their own work. You can find me at dracopol.com.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.