Evaluating the AI-Assisted Developer Experience
This article is a copy-edited transcript of the presentation given at DevFest Melbourne, October 3, 2026. Section titles have been added to help parse the written format. Speaker commentary is in italics. All images in t
This article is a copy-edited transcript of the presentation given at DevFest Melbourne, October 3, 2026. Section titles have been added to help parse the written format. Speaker commentary is in italics. All images in this presentation were generated using Nano Banana Pro.
Hi! I'm Katie, and this is "Evaluating the AI-Assisted Developer Experience".
I really like talks where they give you the summary in the second slide, so here it is:
We burn tokens so you don't have to.
Hopefully the rest of this presentation will make this statement make sense.
So, to help me gauge who we have in the audience, quick show of hands:
- Who here knows what "Agent Skills" are?
- Most of the audience raised their hands.
- Keep your hands up if you have used Agent Skills?
- Somehow, more of the audience raised their hands.
- And keep your hands up if you know that your use of Agent Skills is actually helping your development experience?
- Most of the audience lowered their hands.
Interesting! To help people catch up, here is a refresher:
Agent Skills
From the definition on agentskills.io:
Agent Skills are a lightweight, open format for extending AI agent capabilities with specialized knowledge and workflows.
This is an evolving area of development, but generally: They're a folder of stuff, with a SKILL.md file, and optionally a bunch of other stuff: scripts, references, assets, or other things that help an AI agent complete the instructions. The metadata in the skill includes the name, and description of what the skill does.
Skills can be anything that could help an agent do a task.
You can create your own personal skills that help you with your common workflows, like helping you build a monthly newsletter, or categorize your unread email. You can have agents that sit in your workspace that knows to apply your project's code linting standards after making code suggestions.
Or, you can import skills that have been created earlier.
Google Skills
Google has bunched a bunch of skills!
You can get them at the skills repo: github.com/google/skills
There's around 150 skills and growing.
If you're using any google products, you can install the skill specifically designed to know about that product. There's even a skill in this repo to help find the right skill to use, if you have many installed.
Installing these skills help you with your development needs.
Right?
Do skills actually help?
Are you getting faster, more efficient answers if you use skills?
Because sure, you might hear that people *feel* like skills help. And if you install some yourself, you might think you're getting better results. If you're going to be sharing your skills for others to use, like we're doing, you need to know they will improve things. You need empirical data, just like how you'd test your software before releasing it.
In our case, we can lean on some of the early neural network research in this matter.
Keith-Magee, Russell. 2001. βLearning and Development in Kohonen-style Self Organising Mapsβ, Curtin University. https://hdl.handle.net/20.500.11937/162
AI research has been talking about syllabus presentation and the importance of the order of training for 25 years. When teaching a model, you start with the foundations, then get more refined. The order in which you present information to a student matters.
Since in this case, when you use an agent you're using something that's already baked, you provide the last minute user input or system prompt, that's the thing that takes the most precedence.
Skills are a cheat sheet.
Agent skills are a cheat sheet for agents. Instead of having an agent always have to call out to a web search tool, you can provide the answers to the agent in the form of a skill.
Remember back to that time you had to take an exam, and it was "open book"? You could take the textbook with you! You didn't have to learn what was in the textbook, right? Because you could bring it into the exam!
But when you have the textbook there, and you don't know where the valuable information is inside it, you take forever trying to rifle through it to find the answer, and you end up spending more time answering the question, and when you only have an an hour in an exam setting, you have to have some sort of efficiency in the speed in which you're answering questions, otherwise you're going to run out of time and leave them unanswered. (Not that I know anything about that.)
Agent skills are the one-page and written note you're sometimes allowed to bring in with you. Taking the time to prepare such a sheet is study in itself, and it helps you to know where the useful information is so it's quicker to find and easier to digest, so you can answer questions quickly.
An AI agent will look at what skills are offered based on the skill descriptions, then if that looks like a good candidate, load that skill and all its additional resources, then continue to work out the answer to your question from there. It's using the information provided to it, curated ahead of time, to help answer things quicker than using its baked knowledge corpus, or looking up the entire internet, which may not always yield correct information.
But to test if skills actually make answers faster and more reliable, we need to evaluate the skills are useful to an agent when answering questions.
Do google/skills actually help?
That's the question that is at the heart of the project I've been working on for the last few months.
I've been working on a team that has been evaluating Google Skills, providing ongoing data to skill authors and leadership showing if these skills are actually useful.
We want to prove that if developers install one of our skills, they will have more accurate results, use less tokens, and get faster results.
Evaluation Frameworks
There are a number of evaluation frameworks that you can use, and that list is ever changing.
If you're already in an ecosystem where there is an eval framework, choose that.
DeepEval and harbor are other frameworks that I've heard are useful.
The one we use is Inspect AI. It's developed by the UK AI Security Institute and Meridian Labs, and is used by a number of other AI safety institutions and AI research labs around the globe.
If you're interested in trying out Inspect for yourself, you can try out the codelab at g.dev/ai/eval-aide.
from inspect_ai import Task, task
from inspect_ai.dataset import Sample
from inspect_ai.scorer import model_graded_qa
from inspect_swe import gemini_cli
@task
def skills_eval(skills):
questions = [
Sample(
input="How do I deploy a Cloud Run service?",
target="gcloud run deploy"),
), ...
]
return Task(
model = "google/gemini-3.8-flash"
dataset = questions,
solver = gemini_cli(skills=skills),
scorer = model_graded_qa(),
sandbox = "docker",
)
This is an example eval task, which is written in Python.
This shows a very simple evaluation that includes a number of elements.
- We have the input dataset
- which contains a series of question prompts and target answers
- The solver here is the gemini_cli - the agent that will enact our questions. In this case, we're using the implementation provided by the Inspect SWE package, which contains a number of different software engineering agents.
- And finally the scorer is a "Model-Graded Question and Answer" - an LLM-as-a-judge which checks if the answer contains the target value.
But this is Python. This is a limitation for people who want to use this eval system. Python shouldn't matter.
Language shouldn't matter
"The language of implementation is of no consequence to the end user"
Dr Keith-Magee, the same author of the syllabus order paper from earlier, coined this phrase, and I refer to it often. There's no reason why someone who wants to evaluate their skills needs to also learn Python.
In the case of Google Agent Skills, the people who wrote the skills are googlers with expert knowledge on the particular Google product or feature their skill is about.
They may not know Python, or any programming language. That shouldn't stop them.
YAML Wrapper
We address this by creating a wrapper around inspect AI that allows users to define their jobs in YAML instead of python.
What we're doing here is defining a test matrix.
The test matrix here is defining a combination of a number of different parameters to test against.
The more variations you have, the more different combinations you can get.
In our case, for each prompt:
- We test a number of different models: Models: Gemini (multiple version), and Anthropic's Claude
- We test with "Skill ablation" - how a prompt performs with and without the skill enabled.
- We run the prompt multiple times.
- Inspect AI calls this "epochs".
- We do this because AI is non-deterministic, so if we run the same test multiple times, we can check for consistency.
- We can run the test with different agents:
- Gemini CLI,
- Antigravity CLI,
- Claude Code
- and the Agentic AI Foundation's Goose.
So in the end, we have a configuration that looks like this:
version: 1
description: Weekly Evaluation for Google Skills
models:
- google/gemini-3.8-flash
- google/gemini-3.1-pro
- anthropic/claude-sonnet-5
variants:
- variants.yaml
task_parameters:
dataset: prompts.yaml
eval_tasks.gemini_cli:
agent_framework: "gemini-cli"
include-variants: [baseline, skills_github]
model_roles:
- grader: google/gemini-3.8-flash
We have different models - in this case the most recent flash and pro models (at time of writing), skill ablation, our chosen agent - what runs the task, and our chosen grader: what checks the task.
With this many aspects to our matrix, this makes our tests ... interesting.
Test matrix
For just one sample, you're going to have
- the sample itself,
- but then all the samples you have in your test suite,
- and then you have to test it for every model,
- and then you keep testing it again and again.
The total number of tests you need to run will be
- the amount of samples you have
- multiplied by the conditions you have (with and without the skill)
- multiplied by the number of models you're testing against and then
- multiplied by the epoch to address non-determinism.
This is a Proper Data Science amount of testing we're doing.
Writing Evaluations and Designing Rubrics
But in all of this, you need to actually be writing the correct evaluations.
Much like with concepts of unit testing and code coverage, you can run a lot of tests, but if you're not testing for the right things, your results are going to be poor.
Poor evaluations provide false signals, waste your token budget, and create noise in your metrics.
This part of the presentation then summarised the following DEV posts:
- https://dev.to/googleai/how-to-design-ai-evaluations-you-can-actually-trust-41c3
- https://dev.to/googleai/how-to-write-reliable-rubrics-for-llm-as-a-judge-evaluations-ndp
With all these suggestions, keep in mind:
AI will cheat.
We know that an AI will avoid work if it can help it.
But, with agent skills, we're providing a verified cheat sheet that we want the AI to use. We spent the time writing the one pager of goto information for the exam time, and the AI, when presented with this, will use it.
So if there's best practices, put that in your skill. If there's tools that should be used, add that. Give the agent the best chance of success.
Project Results
So, given all of that, how has our project been going?
We evaluate all skills on a weekly basis. We then process the data which goes into a dashboard showing the skill improvement over time, as skills develop. This has allowed us to give ongoing iterative feedback to authors, without them having to run the evals themselves And we've also spent a lot of time scaling the environment to support the influx of skills published.
Architecture
The architecture is important in this project given the scale we're working with.
Overall the design has been fairly consistent.
We have our eval source code, which has tooling to help us process skills and prompts, which we bundle up and publish to Artifact Registry
Then when we're ready for an evaluation run. The compute pulls the data from artifact registry, and runs the evaluation in sandboxes, which call out to whatever model is needed. Any model hosted there we can support as part of the evals, including Sonnet :)
Inspect AI streams the results to binary .eval files, which we store in Cloud Storage as they are being built, so we can see the progress in Cloud Run, which uses the Inspect AI viewer web service to view the files from Cloud Storage as they come in.
When our evaluation finishes, we write a special file to the storage bucket, which is then picked up for processing. Eventarc triggers when our "done" file is written, which then fires a Cloud workflow that sends the relative information to a Cloud Run job, which performs post processing on the .eval file and stores the results in bigquery, which powers our dashboards.
So most of the aspects of this architecture are product names, but "compute" here seems a bit... nebulous.
Scaling evaluations
We've also had to scale this environment a lot.
If you look at the commit history of the github repo, you can see how many skills have been onboarded over time.
Add our test matrix in, and you can see the numbers we're dealing with.
The slide showed a graph with a line going up to the right, and a number of bar graphs.
You can see the line graph here as number of skills, that is going up and up and up
We published many week on week in July and August, which is when a bunch of this work was rapidly evolving to keep up.
The bars in this graph are the values of our text matrix vectors: Epochs, Models, and Agents.
Except, we can't use a stacked graph here, because all these axes need to be multiplied together
The bar-graph numbers are multiplied together
We're currently running at a total evaluation multiplier of 36, with all the different combinations we have.
Except we're not multiplying that by the number of skills we have.
The line-graph goes from ~150 to ~1100.
We're multiplying that by the number of evaluations we have.
Each skill has two or more evaluation prompts, and each of those are run 36 times.
So the last time I ran the data, we were doing something in the order of 30 thousand different checks?
The spike from July here is when we added too many things at once and had to scale back down again to something sustainable.
Scaling factors
When we first started, we were testing each sample 3 times, but we were asked to increase the epoch to 9, to make sure everything ran the consistently
This was fine, until the amount of tests being performed started ramping up. The spike from earlier was when we added more models and more agents, which meant we were running 50,000 checks, which wasn't something we could reasonably do week on week. One of our vectors had to change to help us support scale.
So, by working with data science experts, we were able to confirm the statistical soundness of reducing our epoch count from 9 to 6. This happened as we were scaling the amount of models we were using, so the total number of evals remained constant (ish).
This was also at a time when we were running these jobs manually, as Engineers in Australia, so we had Monday Business hours to run the evals and have the data ready for American Monday in our dashboards. Given the scale in which we were evolving, it's been ... a while since we've had an eval job complete within business hours.
With this configuration change, we were able to add new models and agents to our test matrix. Previously we were only testing with Gemini models and Gemini CLI, but now we're testing with Anthropic Sonnet and Opus, and using Claude Code as the agent.
Which is important because we want to understand how skills perform when models outside the Google ecosystem are used. We understand that not all developers using Google products use Google's AI tooling to support that, so it's important to understand particularly how the baseline no skill tests work on these models, and how much improvement skills bring.
We tried using Cloud Run jobs to run the evals, which if you add Cloud Scheduler should allow evals to start earlier. However, with this setup we were limited by the available CPUs. Inspect AI by default scales configurations by the amount of available CPUs, so having only base 8 CPU to work with meant our evals did run, but not as fast as we would have liked. So to combat that, we used Cloud Run Jobs' multiple task capability to shard out different chunks of our evals into GKE
The GKE experiments we had hit bottlenecks with the throughput we were able to achieve on this compute platform. The inspect AI ecosystem does have kubernetes sandbox plugins, but at the scale we were running evaluations, we weren't getting the expected results. We suspect we were hitting kubernetes controller limits with the amount of parallel sandbox pods we were scaling up.
So in order to support our growing user requirements, we pivoted to using big ol compute engine instances. With more CPUs, we can run more tasks at once. Also, we can still use Artifact Registry here. We were using Docker containers in our previous setups, but on Compute Engine, we can bundle our code into zip files and publish them to the Generic Artifact Registry!
We've also had to make some adjustments to our evals to support the parallelisation we can afford with this many CPUs available. The bottleneck becomes the Docker network, so adjusting the default subnet there helps scale out.
We're also using Inspect AI's "adaptive connections" that automatically scale up requests on success, and scale down if you're getting too many retries or HTTP 429.
Eval Results
After running all the evals, we get our results!
This is an example view of what Inspect AI gives you after an example run
We can see the matrix here, simplified to just 3.6 and 3.5-flash, with baseline and skills.
And we can compare these values
- While the gemini api baseline example here score 0.28
- the one with skills enabled scored 0.55
- That's a 25% increase!
So it improved!
In the viewer you can click through and see the details about the results, the tool calls, all that.
But that's not exactly accessible to folks that don't dabble in data science.
Eval Dashboard
So, we take the processed data and use it to power a Data Studio dashboard that skill authors can use to more easily see how their skill is tracking.
Here's a view of how the dashboard looks.This is the overview page, where we have some top level averages for all skill ablation
Here's the detail for one skill.
This is an anonymised view of historical data, but it does show a 10 point improvement for cloud run based prompts with the cloud run basics skill enabled.
Scrolling down this page the authors can be linked directly to each prompt and result that sends them back to the Inspect AI viewer to see more detail about the specific eval.
More importantly, they want to see if their skills are green.
If a skill is green then the skill has passed, using an overall pass/fail, based on two factors
- The eval with skill enabled has over 50% accuracy, and more importantly
- The eval with skill enabled improves the no-skill accuracy by more than 5 percent.
- This is important to know that we are getting a big enough change to justify the skill being installed.
The dashboard also has current data points and historical graphs of data over time multiple factors:
- Accuracy: how many rubrics passed
- LLM Turn Count: how many round trips to the model it took to complete the task
- The total working time: how long it took to get the answer, and
- Token efficiency: how many tokens it took to get the answer.
Generally we're looking for a skill having higher accuracy and lower turnout/time/and tokens than its base counterparts.
Summary
So, in summary: We burn tokens so that you have to.
We have months of empirical evidence to prove that if you use agents and work within the Google ecosystem, using the Google Agent Skills will improve your query efficiency and reduce your token usage.
All the resources and references from today's presentation can be found at the QR code.
It's the same QR code from the entire presentation: g.dev/ai/eval-aide
Thank you for your time!
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.










