We made our decision model's API free (and the weights are open)
Most "AI decisions" in production aren't open-ended generation. They're classification in disguise: which team should handle this ticket? is this e-mail phishing? what's the total on this invoice? which sentence in the
Most "AI decisions" in production aren't open-ended generation. They're classification in disguise:
- which team should handle this ticket?
- is this e-mail phishing?
- what's the total on this invoice?
- which sentence in the contract supports that answer?
Teams usually send each of these to a large LLM: 1–2 seconds and a bill per call, a JSON answer to parse, and a "confidence" that means nothing.
We built THX-01 for exactly these decisions. Today we're making its hosted API free.
Try it now (no key, no sign-up)
curl https://api.hal-x.ai/v1/systemone \
-H "Content-Type: application/json" \
-d '{
"state": "Hi, I was charged twice for one order and nobody answers. Refund me!",
"questions": {
"team": {"type": "choice", "question": "Which team handles this?",
"criteria": {"billing": "billing", "tech": "technical", "sales": "sales"}},
"refund": {"type": "noul", "question": "Does the customer want money back?"},
"urgency":{"type": "score", "question": "How urgent?", "criteria": ["low", "medium", "high"]}
}
}'
Every answer comes back with a probability for each option, computed in a single forward pass of about 10 ms.
The free tier allows 200 decisions per minute per IP. One question is one decision.
What it can answer
| type | returns |
|---|---|
choice |
one option, with a probability for every option |
noul |
P(yes) |
score |
a level on an ordinal scale |
number |
a value stated in the document, or null if it isn't there |
excerpt |
a verbatim span with character offsets |
"cite": true |
the passages that support any answer |
How it compares
On our 2,843-ticket benchmark (Azerbaijani, Russian, English, Turkish; clean, corrupted, messy and independently written sets):
| model | avg accuracy | latency |
|---|---|---|
| THX-01 (322M, open) | 98.4 | ~10 ms |
| TypeSafe Jev 1.13 | 97.4 | 331 ms |
| Kev-4B | 92.8 | 830 ms |
The model's confidence is calibrated, so a single threshold lets it handle about 75% of tickets automatically with zero errors on all four test sets.
Run it yourself
The weights are Apache 2.0:
pip install thx01
import thx01
agent = thx01.load("doofz/THX-01")
agent.decide("My card was charged twice", {
"team": {"type": "choice", "question": "Which team?",
"criteria": {"billing": "billing", "tech": "technical"}}})
It runs on a laptop CPU in a fraction of a second.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.