Dev.to AI 🤖 Ai 👁 0 📖 12 min read

When Simpler AI Decisions Worked Better: What I Changed in My Jev-Powered Game Judge

I have been building a small lateral-thinking puzzle game called Sideways. The basic interaction is simple. The player sees a mysterious situation and asks questions until they discover what really happened. For examp

I have been building a small lateral-thinking puzzle game called Sideways.

The basic interaction is simple.

The player sees a mysterious situation and asks questions until they discover what really happened.

For example:

Player:
Is it light?

Game Master:
Yes.

Or:

Player:
Does the wake-up mechanism make sound?

Game Master:
No.

The Game Master is powered by Jev.

I deliberately did not want a general-purpose chatbot generating explanations.

I wanted something much closer to this:

Natural-language input
        ↓
Probabilistic semantic decision
        ↓
Typed result
        ↓
Deterministic TypeScript
        ↓
Game behavior

After several iterations, I realized that this is almost like building probabilistic IF statements over natural language.

And I also learned something I did not expect:

Making those semantic IF statements more sophisticated actually made the game worse.

The fix was not to make the AI smarter.

The fix was to make the decision smaller.

The original idea: AI understands meaning, code controls behavior

The architecture follows a simple rule:

AI = semantic judgment
Code = deterministic policy
User = consequential intent

The model does not directly decide which UI to show.

The browser does not receive the hidden puzzle solution.

The model produces a bounded semantic judgment, and TypeScript turns that judgment into a product result.

For questions, the public result can be:

YES
NO
PARTLY
IRRELEVANT
UNCLEAR

For theory submissions:

SOLVED
PARTIAL
NOT_SOLVED

This separation still feels right to me.

The mistake was not the architecture itself.

The mistake was how much semantic structure I tried to extract from one player sentence.

Version 2: the judge became too clever

In Version 2, a Question judgment looked roughly like this:

{
  wellFormed,
  relevant,
  evidence: {
    entailed,
    contradicted,
    undetermined
  },
  mixedClaims: {
    supported,
    contradicted
  }
}

Most of those values were probabilities between 0 and 1.

Then TypeScript applied rules such as:

if (wellFormed < 0.8) {
  return "REPHRASE";
}

if (relevant <= 0.2) {
  return "IRRELEVANT";
}

if (relevant < 0.8) {
  return "NOT_ENOUGH_INFORMATION";
}

if (
  mixedClaims.supported >= 0.8 &&
  mixedClaims.contradicted >= 0.8
) {
  return "PARTLY";
}

if (evidence.entailed >= 0.8) {
  return "YES";
}

if (evidence.contradicted >= 0.8) {
  return "NO";
}

The Solution Judge was even more detailed:

{
  coreMechanism,
  causalRelation,
  supportingInsight,
  contextualError,
  competingMechanism
}

Again, deterministic TypeScript combined these scores.

The goal was reasonable.

I wanted to distinguish:

"There is a light."
→ PARTIAL

from:

"The light wakes the guest."
→ SOLVED

and from:

"There is a light, but vibration wakes the guest."
→ NOT_SOLVED

Conceptually, Version 2 was much more expressive.

The code also looked disciplined.

And the software tests looked excellent

Version 2 passed:

880 unit/API tests
82 desktop/mobile E2E tests
Leak scanning
Lint
Type checking
Production build

I had calibration data.

I had a separate engineering-generalization corpus.

Hidden grading metadata stayed server-side.

Provider calls remained bounded.

Everything looked good.

Then I used the real model.

The real game felt worse

One puzzle is about a guest waking without sound or physical contact.

The hidden mechanism is a visual light signal.

I tried this:

is it light?

Version 2:

Not enough information

Then:

is it light that wakes the guest?

Version 2:

Not enough information

I submitted the exact idea as a theory:

is it light that wakes the guest?

Version 2:

Not quite yet.

That was already worrying.

Then I tried:

does the wake up mechanism make sound?

Result:

Please rephrase

And even:

does the wake up mechines make sound?

produced:

Please rephrase

The word mechines is obviously a typo.

But the intended meaning is still easy for a human to recover.

At this point the architecture was technically clean but the actual Game Master felt strangely rigid.

What may have gone wrong

There was probably no single cause.

In fact, Version 3 changed several things at once, so I cannot claim that any one of them independently fixed the problem.

But four design problems became obvious.

1. The 0.8 thresholds may have been too conservative

This was my first suspicion.

Suppose Jev had actually returned something like:

wellFormed = 0.74
relevant = 0.92
contradicted = 0.89

for:

does the wake up mechines make sound?

That would mean the system understood quite a lot.

But my code did this first:

if (wellFormed < 0.8) {
  return "REPHRASE";
}

So all the useful information below that gate became irrelevant.

The probability was continuous.

My application policy turned it into a cliff.

0.7999 → rejected
0.8000 → accepted

A lower threshold might have improved the result.

I still think this is a plausible explanation.

But Version 2 had several thresholds, so lowering one would not necessarily solve the whole problem.

2. Several uncertain decisions were connected with hard gates

The deeper problem was not just that 0.8 might have been too high.

There were several such boundaries.

The system effectively became:

IF recoverable enough
AND relevant enough
AND supported enough
AND not contradicted enough
AND ...
THEN YES

Each semantic dimension could be individually reasonable.

But the final behavior depended on all of them interacting correctly.

I had replaced one uncertain AI decision with several uncertain AI decisions and connected them using deterministic cliffs.

That made the overall product more brittle.

3. I decomposed things that humans understand together

Consider:

does the wake up mechines make sound?

A human probably does not consciously evaluate:

grammar quality
→ relevance
→ referent resolution
→ semantic truth

as separate numerical decisions.

We do something more like:

"mechines" probably means "mechanism"
        ↓
They mean the wake-up mechanism
        ↓
They are asking whether it makes sound
        ↓
No

It is one semantic interpretation.

Version 2 decomposed that interpretation into multiple scores.

That was elegant from a TypeScript perspective.

It may have been unnatural from a semantic-decision perspective.

4. The judge knew the scenario, but not explicitly what mystery was being solved

This became one of the most useful changes.

Take:

is it light?

As isolated English, it is ambiguous.

But this is not isolated English.

The player is solving a specific mystery.

For the quiet-alarm puzzle, the actual question being investigated is approximately:

What causes the sleeping guest to wake
when the alarm makes no sound
and nothing touches the guest?

A human Game Master naturally interprets:

is it light?

as:

Is light what causes the guest to wake?

Version 2 had the scenario and hidden answer, but it did not have an explicit representation of the mystery focus.

That changed in Version 3.

Version 3: one semantic decision per judge

Instead of trying to improve Version 2 by adding more examples, more thresholds, or more scores, I removed complexity.

The Question Judge now makes one categorical semantic decision.

Conceptually:

type QuestionSemanticClass =
  | "SUPPORTED"
  | "CONTRADICTED"
  | "MIXED"
  | "UNKNOWN"
  | "IRRELEVANT"
  | "AMBIGUOUS";

Jev returns one probability distribution over those choices.

Then TypeScript performs a deterministic mapping:

SUPPORTED
→ YES

CONTRADICTED
→ NO

MIXED
→ PARTLY

UNKNOWN
→ NOT ENOUGH INFORMATION

IRRELEVANT
→ NOT RELEVANT

AMBIGUOUS
→ PLEASE REPHRASE

There is no active:

wellFormed >= 0.8

gate anymore.

There is no chain of independent semantic Nouls.

There is one semantic classification.

The Solution Judge was simplified in the same way

Version 2 had five independent signals.

Version 3 has one classification:

type SolutionSemanticClass =
  | "CORE_CAUSAL_EXPLANATION"
  | "CORE_WITH_CONTEXT_ERROR"
  | "SUPPORTING_INSIGHT_ONLY"
  | "COMPETING_WRONG_MECHANISM"
  | "NO_MEANINGFUL_INSIGHT";

The product mapping remains deterministic:

CORE_CAUSAL_EXPLANATION
→ SOLVED

CORE_WITH_CONTEXT_ERROR
→ PARTIAL

SUPPORTING_INSIGHT_ONLY
→ PARTIAL

COMPETING_WRONG_MECHANISM
→ NOT_SOLVED

NO_MEANINGFUL_INSIGHT
→ NOT_SOLVED

So I did not abandon semantic distinctions.

I reduced the number of separate decisions required to produce them.

I also added mysteryFocus

Every puzzle now includes a small server-only description of what the player is trying to explain.

For example:

What causes the sleeping guest to wake
when the alarm makes no sound
and nothing touches the guest?

This does not contain the answer.

It only gives the judge the same contextual frame that a human Game Master naturally has.

That helps with short language such as:

is it light?
does it buzz?
is it moving?

without creating hard-coded phrase exceptions.

I shortened the shared rubric

Another change was less visible but important.

Earlier rubrics had gradually accumulated semantic instructions and special distinctions.

Version 3 intentionally moved back toward a smaller question:

Which semantic category best describes this player's statement in this puzzle?

The rubric still tells the model to tolerate:

non-native English
minor spelling mistakes
missing articles
telegraphic wording
ordinary shorthand
natural pronoun resolution

But it does not ask for several independent measurements of those properties.

Again:

Decide less.

What happened after the simplification?

The difference was immediate.

Here are real protected-preview results.

Before: Version 2

is it light?
→ Not enough information
is it light that wakes the guest?
→ Not enough information
does the wake up mechanism make sound?
→ Please rephrase
does the wake up mechines make sound?
→ Please rephrase
is it light that wakes the guest?
[Theory]
→ Not quite yet.

Now compare that with Version 3.

After: Version 3

is it light?
→ Yes
is brightness involved?
→ Yes
something bright wake him?
→ Yes

This is especially important.

The grammar is poor:

something bright wake him?

But the intended meaning is recoverable.

The new judge treated it that way.

Then:

does the wake up mechines make sound?
→ No

The typo no longer caused the system to reject the whole question.

Similarly:

does it buzz?
→ No

and:

is there vibration?
→ No

Short questions also became usable.

Mixed statements started behaving correctly too

I tried:

something bright wake him?
does the wake up mechines make sound?

The result was:

Partly

with the UI:

Some of that is right, but another part isn't.

This is exactly what PARTLY was intended to mean.

Not uncertainty.

Not closeness.

Actual mixed semantic truth.

Most importantly, concise correct theories started solving the puzzle

Question:

does light wake him?

Result:

Yes

Then I reused the exact same text as a theory:

does light wake him?

Result:

You got it.

Another phrasing:

the room gets bright and that wakes him

Question:

Yes

Theory:

You got it.

That is much closer to how a human Game Master should behave.

Wrong causal mechanisms still failed

Simplification did not mean blindly accepting anything containing the correct keyword.

For example:

lamp is there but vibration wakes him

Question:

No

Theory:

Not quite yet.

That matters because false SOLVED results are much worse for this game than conservative partial judgments.

So far, Version 3 became more tolerant without immediately destroying the distinction between correct and incorrect causal explanations.

The behavior also generalized to another puzzle

I tested a different puzzle involving an apparently empty frame whose contents seem to change.

Some examples:

is the frame big
→ Not relevant
is it an animal
→ No
is it a TV
→ No
is it a machine
→ No

Then:

it changes its color with sun light
→ Yes

Submitted as a theory:

it changes its color with sun light
→ You got it.

Again, the grammar is not perfect.

But the causal idea is correct.

And the judge accepted it.

Version 3 is not perfect

There is still some instability around boundary cases.

For example, during one session:

is it the sunrise and sunset
→ Yes

while similar wording later produced:

is it sunrise and sunset
→ Not enough information

That tells me the system is not magically deterministic at the semantic level.

Nor should I expect it to be.

The important difference is that the major player-facing failures changed from:

"I clearly understand what you're asking,
but please rephrase."

to occasional disagreement around genuinely fuzzy boundaries.

For a proof-of-concept game, that is a much better failure mode.

So what actually fixed it?

The honest answer is:

I do not know which individual change fixed it.

Version 3 changed several things together:

Removed multiple continuous semantic gates

Removed the 0.8/0.2 decision chain

Replaced many Nouls with one Choice distribution

Added mysteryFocus

Shortened and generalized the rubric

Made language recovery part of the single semantic classification

Any one of those may have contributed.

The improvement is evidence that the overall architecture became better.

It is not a controlled experiment proving that 0.8 alone was the problem.

A useful future experiment would compare Version 2 under several threshold policies.

That could answer:

Was Jev already producing useful semantic probabilities that my application was throwing away?

I think that is entirely possible.

Jev started feeling less like "adding AI" and more like building probabilistic IF statements

This project changed how I think about decision models.

I am not really building a chatbot.

I am building something conceptually closer to:

if (semanticCondition) {
  doSomething();
}

Except the condition is not:

x > 10

It is:

Does this player's sentence semantically express
a proposition supported by the hidden explanation?

Jev estimates that semantic condition.

TypeScript decides what happens next.

That is why I increasingly think of this pattern as:

probabilistic IF statements over natural language

or:

an AI-powered switch/case

Version 2 accidentally turned that simple idea into a giant conditional expression.

Version 3 moved it back toward a small semantic switch.

There is also a hidden engineering cost: calibration

Another lesson was that API pricing is not the only cost of using a decision model.

For this project I needed to repeatedly:

design semantic categories
test real player wording
test spelling mistakes
test short questions
test false mechanisms
test incomplete solutions
check false positives
check false negatives
change policies
retest on protected Preview

That is human labor.

Jev reduced one kind of complexity for me.

I did not have to parse arbitrary assistant prose into application state.

But that complexity did not disappear.

Some of it moved into:

decision design
calibration
threshold selection
test-play
behavioral evaluation

I would therefore not say:

Jev requires more engineering work than a traditional LLM.

I have not run a controlled comparison that proves that.

A more accurate statement is:

In this project, Jev moved a meaningful part of the engineering effort from output handling to decision design and calibration.

That is an important cost to account for.

849 tests still did not replace test-play

Version 3 currently passes:

849 unit/API tests
82 desktop/mobile E2E tests
Leak scanning
Lint
Type checking
Build

Those tests are valuable.

They prove things such as:

the parser rejects invalid distributions
the deterministic mapping works
private grading data does not reach the browser
provider call counts remain bounded
the UI still behaves correctly

But they still cannot fully prove:

A human writes an unseen sentence
        ↓
The model understands it as intended

That requires actual behavioral testing.

For AI systems I now think about validation as two separate layers:

Layer 1
Software correctness

Layer 2
Model behavior

A green CI pipeline is excellent evidence for Layer 1.

It is not sufficient evidence for Layer 2.

Why I am stopping here

There are still things I could tune.

I could add more categories.

I could add confidence gates.

I could add more examples.

I could create exceptions for phrases like:

sunrise and sunset

I am deliberately not doing that.

That path is exactly how Version 2 became complicated.

For this PoC, Version 3 is now a strong freeze candidate.

The remaining requirement is not:

perfect semantic classification

It is:

Natural short player language usually works.

Minor grammar and spelling errors usually work.

Correct causal explanations can solve the puzzle.

Wrong causal explanations do not solve it.

Hidden information stays hidden.

The game remains fun.

That is enough.

What if this still breaks later?

I already have a fallback.

Simplify even further:

YES
OTHER

OTHER would intentionally collapse:

NO
UNKNOWN
IRRELEVANT
AMBIGUOUS
possibly MIXED

That would lose information.

But if it produced a better game experience, I would seriously consider it.

Again:

The goal is not to extract the maximum amount of semantic information from the model.

The goal is to ask for the smallest useful decision.

Final takeaway

The most surprising lesson from building Sideways has been this:

More semantic structure does not automatically produce better AI behavior.

Version 2 looked more sophisticated.

It had more probabilities.

More semantic axes.

More explicit thresholds.

More detailed grading logic.

And worse real gameplay.

Version 3 asked Jev to do less:

One semantic Choice
        ↓
One probability distribution
        ↓
Deterministic TypeScript switch

And the Game Master got noticeably better.

So the principle I am taking forward is:

Do not ask AI to make a more complicated decision than your product actually needs.

Or even shorter:

Understand enough.

Decide less.

That turned out to be a much better architecture for this game.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.