Dev.to AI 🤖 Ai 👁 0 📖 2 min read

We made our decision model's API free (and the weights are open)

Most "AI decisions" in production aren't open-ended generation. They're classification in disguise: which team should handle this ticket? is this e-mail phishing? what's the total on this invoice? which sentence in the

Most "AI decisions" in production aren't open-ended generation. They're classification in disguise:

  • which team should handle this ticket?
  • is this e-mail phishing?
  • what's the total on this invoice?
  • which sentence in the contract supports that answer?

Teams usually send each of these to a large LLM: 1–2 seconds and a bill per call, a JSON answer to parse, and a "confidence" that means nothing.

We built THX-01 for exactly these decisions. Today we're making its hosted API free.

Try it now (no key, no sign-up)

curl https://api.hal-x.ai/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{
    "state": "Hi, I was charged twice for one order and nobody answers. Refund me!",
    "questions": {
      "team":   {"type": "choice", "question": "Which team handles this?",
                 "criteria": {"billing": "billing", "tech": "technical", "sales": "sales"}},
      "refund": {"type": "noul", "question": "Does the customer want money back?"},
      "urgency":{"type": "score", "question": "How urgent?", "criteria": ["low", "medium", "high"]}
    }
  }'

Every answer comes back with a probability for each option, computed in a single forward pass of about 10 ms.

The free tier allows 200 decisions per minute per IP. One question is one decision.

What it can answer

type returns
choice one option, with a probability for every option
noul P(yes)
score a level on an ordinal scale
number a value stated in the document, or null if it isn't there
excerpt a verbatim span with character offsets
"cite": true the passages that support any answer

How it compares

On our 2,843-ticket benchmark (Azerbaijani, Russian, English, Turkish; clean, corrupted, messy and independently written sets):

model avg accuracy latency
THX-01 (322M, open) 98.4 ~10 ms
TypeSafe Jev 1.13 97.4 331 ms
Kev-4B 92.8 830 ms

The model's confidence is calibrated, so a single threshold lets it handle about 75% of tickets automatically with zero errors on all four test sets.

Run it yourself

The weights are Apache 2.0:

pip install thx01
import thx01
agent = thx01.load("doofz/THX-01")
agent.decide("My card was charged twice", {
    "team": {"type": "choice", "question": "Which team?",
             "criteria": {"billing": "billing", "tech": "technical"}}})

It runs on a laptop CPU in a fraction of a second.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.