Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 8 min read

Jev Gives You Probabilities, Not Decisions

In my previous post, I showed how to use Jev in cloud engineering scenarios: filtering NSG rules, checking Bicep what-if output, and linting Azure Policy. I didn't explain how it works. This post covers the concepts and

In my previous post, I showed how to use Jev in cloud engineering scenarios: filtering NSG rules, checking Bicep what-if output, and linting Azure Policy. I didn't explain how it works. This post covers the concepts and the mental model behind Jev. The examples still use PowerShell and Azure, but the ideas apply to any language and any scenario.

I use Doug Finke's Jev PowerShell module, but you can use any language.

The principle

TypeSafe, the company behind Jev, was inspired by Thinking, Fast and Slow, the bestseller everyone pretends to have read, by psychologist Daniel Kahneman. Kahneman describes two systems of thought. System 1 is fast and instinctive, like answering 2 + 2. System 2 is slow, effortful and logical, like counting how many times the letter a appears on a page.

Jev is built on the System 1 idea. It's a new class of AI model: it doesn't generate text, it makes fast, structured decisions. An LLM takes text and returns text that you then have to parse. Jev takes text and returns a typed decision with probabilities and a confidence score, which you can use directly as a signal in a script, a program or a workflow.

The cost model is different too. You only pay for input: $42 per billion input tokens, or $0.042 per million. Output is free, because it's always small. TypeSafe also claims response times between 70 and 500 ms (you need to add the network round trip). It seems coherent during my testing.

Any Jev call use two inputs, the state or the context to evaluate and questions about the state.

State: the context

The state is the context you want to evaluate using Jev. It can be a message, a log, or application output.

Jev supports three formats: string, JSON object, or array.

`String: The customer is blocked by a connection issue to the database.

JSON Object

json
{
"securityRules": [
{
"name": "allow-https",
"properties": {
"description": "Public HTTPS for the web site. Owner: web team",
"protocol": "Tcp", "sourcePortRange": "*", "destinationPortRange": "443",
"sourceAddressPrefix": "Internet", "destinationAddressPrefix": "*",
"access": "Allow", "priority": 100, "direction": "Inbound",
"sourcePortRanges": [], "destinationPortRanges": [], "sourceAddressPrefixes": [], "destinationAddressPrefixes": []
}
}
}

Array:

json
["Inbound", "Allow", "HTTPS"]

TypeSafe recommand to use a JSON object, if possible and the string for simple task.

The state contains the data, never the question. Clean it up when you can: remove empty fields, IDs and timestamps that don't matter for the decision. You save tokens and the model gets a clearer signal. The context window is limited to 32k tokens, and non-English texts gives less accurate results.

Questions: the three primitives

A Jev request evaluates the state against one or more questions. Each question is evaluated independently.

There are 3 types of question: Score, Noul and Choice.

Score

A Score question rates the state against an ordered list of 2 to 10 descriptive levels, from lowest to highest. These levels are the criteria

powershell
$question = New-JevQuestion -Name risk -Type Score

-Instructions 'How risky is it to deploy the changes in whatif?'
-Criteria @(
'No risk: cosmetic or read-only property changes'
'Low: new isolated resources'
'Medium: changes to shared network resources (VNets, peerings, route tables)'
'High: an outage or a security incident is likely'
)
`

To get the result:

`powershell
$result = Invoke-Jev -State $summary -Question $question
$result.answers.risk | convertTo-Json -Compress -Depth 10
`

The risk object look like

`json
{
"type":"score",
"score":2.49,
"confidence":0.5,
"legend":{
"0":"No risk: cosmetic or read-only property changes",
"1":"Low: new isolated resources",
"2":"Medium: changes to shared network resources (VNets, peerings, route tables)",
"3":"High: an outage or a security incident is likely"
},
"probabilities":{
"0":0.0,
"1":0.01,
"2":0.5,
"3":0.49
}
}
`

  • The type, is the type of the question.
  • The score is the probability-weighted mean of the level, here 2.49, between Medium and High.
  • Legend, each level and their number.
  • Probability, Probability for each level, the sum of each probability is equal to 1.
  • Confidence, a number between 0 and 1 that shows how the probabilities are concentrated.

In the example the probability for High and Medium are similar, and the confidence is 0.5. It's a coin flip between Medium (2) and High (3), the model can not judge, an humain should take a look. It is not "medium-high" risk.

Noul

A Noul question asks a yes/no question and returns the probability that the answer is yes. The criteria are optional, but if you set them, you must describe both true and false.

`powershell
$feedback = [pscustomobject] @{
message = 'The customer is blocked by a connection issue to the database.'
}

define the criteria for the question

$question = New-JevQuestion -Name noulquestion -Type Noul
-Instructions 'Does the report give steps that would reproduce a defect?'

-Criteria @{
true = 'The report describes a defect and gives concrete steps to reproduce it.'
false = 'The report does not give concrete steps to reproduce a defect.'
}

invoke the question with the current state

$result = Invoke-Jev -State $feedback -Question $question
$result.answers.noulquestion
`

The result:

`json
{
"type":"noul
noul":0.05
}
`

  • Type of question; noul
  • The noul property, the probability that the answer is true

The noul property will give you the probability that the answer to the question is true. The probability of false is simply 1 βˆ’ noul.

Choice

A Choice question selects one option from a set of 2 to 255. The criteria are a hashtable: each key is an option name and each value describes it.

The instruction is the question for the model.
And the criteria is an array of map. Add an unclear option to give the model a way out. Without one, it has to pick a team even when none fits.

`powershell
$criteria = @{
support = 'The issue needs technical support.'
sales = 'The issue concerns sales.'
unclear = 'It is unclear which team should handle the issue.'
}

create the question based on the criteria

$question = New-JevQuestion -Name TeamRouting -Type Choice -Instructions 'Which team should handle this?' -Criteria $criteria

invoke the question with the current state

$result = Invoke-Jev -State $feedback -Question $question

check the result of the question

$result.answers.TeamRouting
`

The result:

`json
{
"type":"choice",
"choice":"support",
"confidence":1.0,
"probabilities":{
"support":1.0,
"unclear":0.0,
"sales":0.0
}
}
`

  • The type, is the type of the question
  • Choice gives you the slected option.
  • Confidence gives you a number between 0 and 1 on how probability is spread on the different options. A 1.0 means all the probability is on one option; lower values mean it's spread across several.
  • Probabilities for the different option, the sum is always equal to 1.

Probability, confidence and calibration

Jev doesn't really give you decisions. It gives you probabilities, it is your role to turn them into decisions.

For each question type you have probabilities, and you have confidence. Probabilities tell how likely the model thinks each option or level is. The distribution and the concentration between the possible answers is also important. This is how confidence is calculated.

If the probabilities are concentrated on one option. For example a=0.9, b=0.06 and c=0.04. All the probability are on the option a, the model is confident.

If they're spread across several options, a=0.4, b=0.33 and c=0.27. There is no clear winner, the model is uncertain.

The noul question type doesn’t return a confidence, just the probability that of the true answer. But you can calculate it via this formula:

`Confidence = | 2p – 1|

Where p is the probability of the yes answer

  • If p = 0.05 the confidence will be 0.9. Meaning the response is false with a large confidence
  • If p = 0.95 the confidence will be 0.9, the confidence for the true is large.
  • If p = 0.4 the confidence will be 0.2, the can choose an option and an humain should be involced.

The answer to the question gives you β€œThe What”, the confidence tells you how to act with it.

If the confidence is higher or equal to 0.9 you should automated it, between 0.9 and 0.5 it needs a confirmation and bellow 0.5 it new a human investigation.

This is the recommendations of TypeSafe. But you should build your own thresholds with your data.

You need to take a set of data and create 30 to 60 different states where you know the correct answer and run tests.

The average confidence of correct answers doesn't give you a threshold. Instead, sort results by confidence and find the lowest confidence above which accuracy meets your target (85% for example). That's your automation threshold.

Writing good questions

There is another way, to get the most of Jev is to know how to make good question. Question are the equivalent of a prompt in the LLM world.

There are several patterns:

Question should be atomic. Only one concern per question. For example, "Is the rule risk and undocumented" should be two questions: "is the rule risky" and "is the rule undocumented".

Do not use Jev to get an answer a script can easly have. If it is deterministic, keep it in the code. Is port 22 open to the Internet?" is a script check. "Does this rule's description identify an owner?" is a Jev question.

Criteria is the real prompt, the state is the data, Avoid vague and ambiguous criteria. Take a look at these options: "High: dangerous" versus "High: an outage or a security incident" is likely too ambigous. For Choice, options must not overlap (Ex "Tech Support" and "IT Support").

Batch question in one call, Jev can do parallel checks

Example:

`powershell
$questions = @(
New-JevYesNoQuestion -Name security
-Question 'Do the changes in
whatif weaken the security of the environment?'
-TrueCriteria 'Password login enabled on a VM, no NSG on a NIC or subnet, a security setting removed or relaxed' `
-FalseCriteria 'Security settings are unchanged or stronger'

New-JevYesNoQuestion -Name public_ip `
    -Question 'Do the changes in `whatif` expose a resource to the internet with a public IP?' `
    -TrueCriteria 'A public IP address is created or attached to a NIC, a VM or a load balancer' `
    -FalseCriteria 'No public IP is created or attached'

New-JevQuestion -Name risk -Type Score `
    -Instructions 'How risky is it to deploy the changes in `whatif`?' `
    -Criteria @(
        'No risk: cosmetic or read-only property changes'
        'Low: new isolated resources'
        'Medium: changes to shared network resources (VNets, peerings, route tables)'
        'High: an outage or a security incident is likely'
    )

)

$state = @{ whatif = $summary -join "`n" }
$answers = (Invoke-Jev -State $state -Question $questions -Mock:$Mock -Raw).answers




## conclusion

Despite its name, what a decision model really gives you is probabilities: a signal that tells you when to automate and when to ask for review.

Remember four things. The state is data, never the question. Each question covers one concern. The criteria are the real prompt. Measure your own thresholds, especially if your data isn't in English or state is very specific.

Keep the limits in mind too. Jev returns numbers without reasoning, so you can't audit why it decided.
πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.