Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 7 min read

A Simple Text Skill to Address AI Arithmetic Errors

Can an ordinary free chat handle a task that requires both understanding the problem and working carefully through a sequence of calculations? That was the practical question I wanted to explore. We deliberately chose GP

A Simple Text Skill to Address AI Arithmetic Errors

Can an ordinary free chat handle a task that requires both understanding the problem and working carefully through a sequence of calculations? That was the practical question I wanted to explore. We deliberately chose GPT-5.6 Chat Free to see how far we could get with serious numerical work in a mode available to many users. This was not about saving money on a subscription. If an approach helps in everyday chat, it is more useful than one that works only in a specially configured environment.

We used a packaging machine selection task. It included input data for several options, calculation rules, and a final comparison. The specific domain is not essential here: a similar error could occur in a cost estimate, a financial spreadsheet, an application ranking, or any other lengthy numerical answer. A model may understand the problem correctly and still make a mistake in a single arithmetic operation. If that number then feeds into the final result, a convincing answer becomes an incorrect one.

An arithmetic error in the final table

In one run without the skill, GPT-5.6 Chat Free made a clear arithmetic error.

The screenshot below shows the row for one option, UPK-0130. The model itself reported three intermediate values for it: 0.8565, 0.5000, and 1.0000.

Just above the table, it had stated the scoring rule: 0.60 Γ— the first value + 0.10 Γ— the second + 0.30 Γ— the third. Substituting the numbers gives:

0.60 Γ— 0.8565 + 0.10 Γ— 0.5000 + 0.30 Γ— 1.0000 = 0.8639

The table, however, shows 0.8279.

The difference is 0.0360: a direct arithmetic error that can be checked using the model's own numbers.

The UPK-0130 row in the run without the skill, with the reported score of 0.8279 underlined in red

Screenshot 1. Run without the skill: the UPK-0130 row with a score of 0.8279 is underlined in red.

We noticed this error by chance, only because we were following UPK-0130's position in the ranking and checked that particular row. We did not recalculate the entire table. The same kind of arithmetic error could have occurred anywhere else without attracting our attention. A user who does not check the model's numbers with a calculator would probably miss it.

In a study by Jacob Beck and colleagues, 2,784 participants checked values from report tables presented to them as AI answers. The authors write that β€œeven with pre-annotations and human adjudication, substantive errors can go unnoticed.”

The answer looks convincing: a detailed table, a seemingly correct formula, four decimal places, and even a section on result validation. But careful formatting does not replace multiplication and addition. In this run, the error lowered UPK-0130's score and distorted the final comparison table. That gave us a starting point: we needed a rule that treated calculation as a separate, verifiable task, rather than another polished answer template.

How researchers separate reasoning from computation

We did not come up with the underlying idea from scratch. In Program-Aided Language Models (PAL), researchers propose separating two activities: the language model interprets the problem and writes the solution steps, while a conventional interpreter performs the arithmetic. Program of Thoughts develops a similar approach: the model handles reasoning and delegates computation to a tool. Calc-X and Calcformers explore the same idea on arithmetic tasks with access to a calculator.

For me, the practical lesson is straightforward: a model is useful as an assistant that understands text and builds a procedure, but important numbers are better calculated by a mechanism whose results can be reproduced.

We are exploring how far we can get with numerical work using widely accessible tools, such as GPT-5.6 Chat Free. That is why our experimental Deterministic Calculation Skill v0.1 takes the form of a single text file, SKILL.md. In an ordinary chat, this file acts as a prompt. In our test, we downloaded it from GitHub, attached it to the chat, and first sent a separate setup prompt. We submitted the calculation task in the next message.

Here is the exact setup prompt:

For all calculations in this chat, use the attached DETERMINISTIC_CALCULATION_SKILL as a mandatory local instruction. Computed numerical values must not be generated by the language model itself. If a deterministic calculation tool is available, use it. If such a tool is unavailable and the calculation cannot be performed reliably and deterministically, stop the calculation and report that limitation. Do not solve the task yet.

In Codex, the same file can be installed from GitHub as a skill. It then becomes part of the available skill set: Codex sees the skills' names and descriptions and can load the relevant skill when a task calls for it. Our skill contains only text instructions; other skills can also include scripts, reference material, and templates.

What our skill does

The skill does not contain ready-made solutions to a particular task. It defines a procedure for working with numbers:

  • Clarify the calculation rules before starting. If an important rule is missing, ask the user instead of inventing one.
  • Obtain calculated values using an available calculator, code, spreadsheet, or another reliable tool.
  • Preserve the original numbers, units, and precision without silently replacing or β€œcorrecting” the data.
  • Show the formula, input values, important intermediate steps, and final result so that the calculation can be repeated.
  • Keep source data, transformed values, and final metrics distinct.
  • Validate the result after calculation: numerical ordering, equalities, bounds, signs, and possible double counting.
  • If no reliable calculation method is available, say so explicitly instead of presenting a guess as a result.

This gives each number in the answer a traceable origin that can be checked.

What changed in the run with the skill

We continued testing with the Deterministic Calculation Skill. The second screenshot shows part of GPT-5.6 Chat Free's answer when using the skill.

It is particularly interesting because the result was not flawless. In the normalization step, the model wrote the fraction:

(7,821.428571 βˆ’ 2,346.428571) / (7,821.428571 βˆ’ 730)

It reported 0.778916 as the result. Evaluating that exact fraction gives approximately 0.772059.

The model then substituted the incorrect 0.778916 into the final score calculation and reported 0.803535. However, the sum using the displayed 0.778916 is approximately 0.805593. Two adjacent lines in the screenshot therefore contradict each other.

Run with the skill: the incorrect intermediate value and the final score of 0.803535 are marked in red

Screenshot 2. Run with the skill: the incorrect intermediate value and the final score of 0.803535 are marked in red.

There is still a key difference from the first run. Using the correct normalization value, approximately 0.772059, in the final calculation produces 0.803535 after rounding. In other words, the final score and the selected option, UPK-0130, were correct, although the model did not correct the intermediate text.

I consider the correct final score a result of using the skill: it required the calculation to be checked, and the model appears to have used the correct value in the final number. It did not show the recalculation itself. The incorrect intermediate lines remained in the answer, but the explicit calculation chain allowed us to spot them.

What the skill and Astra have in common

We also looked at GPT-6 Astra's answers to the same task. I was interested less in the ranking positions than in how Astra described its numerical checks. In two screenshots, it explicitly states that it recalculated the scores and separately checked the ordering of the options.

GPT-6 Astra Extra High: the model reports an independent recalculation of the scores and pairwise checks of the option ordering

Screenshot 3. GPT-6 Astra Extra High: the model reports an independent recalculation of the scores and pairwise checks of the option ordering.

Another run, with GPT-6 Astra Max, contains almost the same statement: the scores were independently recalculated, and the final order was checked against every pairwise comparison.

GPT-6 Astra Max: the end of the answer also reports recalculated scores and pairwise checks of the final order

Screenshot 4. GPT-6 Astra Max: the end of the answer also reports recalculated scores and pairwise checks of the final order.

There is a clear shared logic here: obtain a numerical result, then check it separately, including the final ordering. Our skill makes that check an explicit rule. It also requires calculated values to come from an available deterministic tool, preserves the source data, and calls for enough visible steps to reproduce the result.

OpenAI's official GPT-6 Astra documentation lists support for Code Interpreter. This tool can execute Python code, including for mathematical tasks. That capability fits the idea of separating arithmetic from writing the answer. However, the screenshots do not establish whether Astra used Code Interpreter in these particular runs: they show its statements about validation, rather than records of tool calls.

For me, that is the methodological similarity. Additional numerical checks are useful beyond a powerful model: an ordinary GPT-5.6 Chat Free can also be asked to perform them. That is why we put the rules into a portable text skill that can easily be attached to a chat prompt.

What this check showed

What matters to me is the practical result: in an ordinary GPT-5.6 Chat Free session, a short instruction established a stricter procedure for working with numbers, and the second run's final result held up under an independent arithmetic check.

Together, these runs show both the usefulness and the limits of the approach. A correct final result does not mean that every intermediate line in the answer is correct. Validation should therefore cover the entire chain, rather than stop at the final number. Astra's example helps show that additional numerical checks are a sensible working practice, rather than an arbitrary requirement of our skill, and that this practice can be made accessible in a simple chat.

Materials

AI disclosure: The ideas, main points, and substantive revisions are the author’s. AI assisted with editing the text.

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.