Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 2 min read

Questions for a chatbot

Today I have a file open with an empty table. At the top are the numbers from two weeks ago: out of 68 answers, one passed. At the bottom, the hypotheses I wrote so I wouldn't cheat myself when measuring again. The table

Today I have a file open with an empty table. At the top are the numbers from two weeks ago: out of 68 answers, one passed. At the bottom, the hypotheses I wrote so I wouldn't cheat myself when measuring again. The table in the middle, the one that would say whether the chatbot got better, has been empty for 16 days.

What I ask it

The chatbot answers questions about pasture growth rates with a chart and a text that explains it. Its user is a rural worker who doesn't code. To find out whether it answers well, I built a golden set: 68 questions, none of them from real users. Twenty came from a first round of examples and 48 from a second.

A golden set is a fixed list of questions with what you expect from each answer, which you rerun after every change to see whether things got better. If whoever grades only says "pass" or "fail", you know how much fails. My auditor, a separate agent, returns something besides the verdict: faltantes (what each answer was missing), with a type from a closed vocabulary (data, domain knowledge, tool, prompt) and a key so they can be summed across questions. That way you know what to fix.

1 of 68

In the September 14 baseline, 1 passed; 31 passed with reservations, 35 were rejected and 1 errored out.

That number wasn't the useful part. The useful part was 205 gaps, which grouped into 102 keys and mapped to 22 tickets: 16 new and 6 that were already in the backlog. Today 11 are done, 5 cancelled and 6 open. The three most frequent gaps: not knowing the dataset's reference date (24), having no rainfall column (19) and not being able to filter by field and paddock at the same time (16).

Rainfall didn't get fixed. There is no weather data, and I decided the chatbot should say so instead of inventing a proxy. A gap that gets resolved by saying "I don't know".

The pending measurement

Before measuring again I wrote hypotheses with a threshold: rejected answers, from 35 down to 20 or fewer; the reference date, from 24 down to 5 or fewer. And a counter-hypothesis: the change that lets the model call no tool at all could make questions that were chartable worse.

The second run generated the answers, but the audit stage, about 28 dollars, was cut off: the API credit ran out. It has been blocked since September 15.

I know what I fixed. I don't know if it got better.

When you evaluate something you built, does it hand you a score or a list of what's missing? And do you write down beforehand what you expect to change?

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.