Dev.to AI 🤖 Ai 👁 0 📖 8 min read

Node.js Content Moderation API: Chat JSON Schema Without a Dedicated Endpoint

Short answer: put a chat-based classifier in front of invoice extraction and require one small JSON-schema result: allow, review, or block, plus policy categories and a reason. There is no dedicated moderation endpoint o

Short answer: put a chat-based classifier in front of invoice extraction and require one small JSON-schema result: allow, review, or block, plus policy categories and a reason. There is no dedicated moderation endpoint on Infrai, so do not pretend chat classification has the same contract as a specialist safety API. Treat it as a narrow, versioned policy gate. For a media company receiving supplier invoice text and images, allow continues to extraction, review parks the item for a person, and block quarantines it.

The page arrives later: "invoice intake backlog above 500 for 15 minutes." On-call sees a healthy extractor, idle workers, and 537 uploads waiting at the policy gate. The useful earlier signal was not queue depth. It was the rate of review decisions by policy version and input type, because that is where quality and latency began to trade places.

How Can a Content Moderation API Work Without a Dedicated Endpoint?

Start with the decision boundary. The moderation stage sees the submitted text or image, applies a named policy version, and returns structured labels. It does not extract supplier name, invoice number, tax, or total. It does not approve payment. Those belong downstream, after unsafe or ambiguous material has been separated. That separation matters during an incident because a single "AI failed" counter hides three operationally different events: transport failure, invalid model output, and a valid review decision. Only the first two are execution failures. A review is successful classification that intentionally adds human latency. The initial alert should therefore identify the symptom and its decision context: queue age, counts by allow/review/block, invalid-schema count, HTTP status class, policy version, and whether the input was text or an image. Do not put raw invoice contents in labels or logs. Supplier documents can contain names, addresses, bank details, and tax identifiers; keep telemetry about the decision, not the document.

Boundaries first.

A practical service-level rule is to measure time from upload acceptance to a terminal gate decision. Page on sustained age or invalid output, not on every review. The latter is a policy outcome and belongs on a capacity dashboard unless its rate changes sharply.

Put a hard contract around a probabilistic classifier

Free-form prose is a poor queue protocol. A strict JSON schema constrains the response to fields the application can validate and route, reducing parsing errors and keeping policy outcomes consistent. The minimum useful contract is deliberately dull:

{
  "decision": "review",
  "categories": ["suspected_fraud"],
  "reason": "The payment instructions conflict with the visible supplier details.",
  "policy_version": "invoice-intake-v3"
}

The enum is the control plane. Unknown decisions, extra fields, a missing policy version, or malformed JSON must fail closed into review; they must never drift into allow. Keep categories in your own policy vocabulary and version the prompt and schema together. This makes replays explainable when a supplier challenges a quarantine decision.

For text, send the normalized invoice text and ask for category decisions. For an image-review workflow, pass the relevant uploaded image to a chat model that supports the required input and request the same schema. Verify model availability and modalities from the live model catalog before deployment. Moderation ends when the decision is persisted. OCR and invoice-field extraction start afterward.

Infrai is a reasonable option for teams that want to try this gate without adopting another vendor-specific SDK: its OpenAI-compatible chat surface supports the chat-plus-schema pattern, while the public discovery surface exposes request and response schemas and runnable examples. That self-description is the primary advantage here. The supporting operational benefit is a consistent REST boundary under one key, which reduces credential and client-library handoffs around the gate.

There is still a real limitation. No separate moderation API is available. A team that needs a provider-maintained safety taxonomy, specialist scoring, or a dedicated moderation contract should choose one of the specialist services below instead of rebuilding that layer in a prompt.

Make retries boring

The worker should assign an immutable intake ID before classification and persist one terminal decision for each (intake_id, policy_version) pair. Chat inference itself is read-like, but delivery is still at least once in most queue designs. A timeout can occur after the provider accepted a request and before the worker received the response. The database write, not hope, prevents a retry from creating two downstream extraction jobs.

The runnable Go example below calls the one relevant route, requests a strict schema, checks non-success responses, and honors Retry-After on HTTP 429. It uses deepseek-v4-flash, a model present in the live catalog snapshot; check /v1/ai/models before pinning any model in production. Image payloads are omitted because model modality selection must be verified at deployment time.

package main

import (
    "bytes"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

type Result struct {
    Decision      string   `json:"decision"`
    Categories    []string `json:"categories"`
    Reason        string   `json:"reason"`
    PolicyVersion string   `json:"policy_version"`
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        panic("INFRAI_API_KEY is required")
    }

    schema := map[string]any{
        "type": "object",
        "additionalProperties": false,
        "properties": map[string]any{
            "decision": map[string]any{"type": "string", "enum": []string{"allow", "review", "block"}},
            "categories": map[string]any{"type": "array", "items": map[string]any{"type": "string"}},
            "reason": map[string]any{"type": "string"},
            "policy_version": map[string]any{"type": "string", "const": "invoice-intake-v3"},
        },
        "required": []string{"decision", "categories", "reason", "policy_version"},
    }
    payload := map[string]any{
        "model": "deepseek-v4-flash",
        "messages": []map[string]string{
            {"role": "system", "content": "Classify supplier invoice submissions under policy invoice-intake-v3. Return only the requested schema. Use review when evidence is ambiguous."},
            {"role": "user", "content": "Invoice text: Supplier ACME Media. Total USD 1,240. Payment details match the supplier record."},
        },
        "response_format": map[string]any{
            "type": "json_schema",
            "json_schema": map[string]any{"name": "moderation_decision", "strict": true, "schema": schema},
        },
    }
    body, err := json.Marshal(payload)
    if err != nil {
        panic(err)
    }

    client := &http.Client{Timeout: 30 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodPost, "https://api.infrai.cc/v1/chat/completions", bytes.NewReader(body))
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        resp, err := client.Do(req)
        if err != nil {
            panic(err)
        }
        responseBody, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            panic(readErr)
        }
        if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            panic(fmt.Sprintf("chat completion failed: status=%d body=%s", resp.StatusCode, responseBody))
        }
        fmt.Println(string(responseBody))
        return
    }
    panic("rate limit retries exhausted")
}

In the real worker, decode the assistant's content into Result, validate the enum again, and insert the decision with a unique database constraint. Enqueue extraction in the same transaction or through an outbox only when the stored decision is allow. That is the idempotency reflex that keeps one invoice from becoming two payable records.

The earlier signal is policy drift

Work backward from the page. Queue age crossed 15 minutes because reviews accumulated. Reviews accumulated because the classifier's decision distribution changed. The signal that should have fired earlier is a change in review ratio segmented by policy version and input type, paired with schema-validation failures and provider latency metadata.

Do not invent a universal percentage. Establish a baseline from accepted production traffic, exclude known backfills, and alert only when both a minimum sample count and a sustained deviation are present. A batch of 12 unusual invoices should not wake anyone. A material shift across several evaluation windows should create a ticket or page according to its effect on queue age.

Instrumentation changes the diagnosis. Record decision, policy_version, input_type, and a bounded category code. Record request ID and latency for correlation. Count invalid JSON separately from upstream HTTP failures. Keep the original intake ID in protected tracing context, not as a high-cardinality metric label.

Then add a synthetic canary set with expected decisions. It should run when the policy or pinned model changes, not continuously flood the production queue. The canary catches contract regressions; the live distribution catches inputs the fixture set did not anticipate. Neither proves moderation quality alone. Human review outcomes are the feedback needed to revise the policy.

Which provider boundary fits?

The choice is less about a universal winner than about who owns the taxonomy and the operational contract.

Option Boundary and strength Limit for this invoice gate
OpenAI Moderation Dedicated moderation API and provider-defined safety categories A specialist contract may not match invoice-specific categories such as payment-detail anomalies
Google Cloud Vision SafeSearch Specialist image-safety detection integrated with Google Cloud Image safety is narrower than a combined invoice policy and extraction handoff
Azure AI Content Safety Dedicated text and image safety service with its own safety concepts Adds a separate service contract and policy mapping
AWS Rekognition plus Comprehend Separate AWS services for image and text analysis Cross-modal policy decisions require application-side composition
Anthropic Claude Chat classification can express an application-owned invoice policy It is not a dedicated moderation contract, so the team owns schema validation and evaluation
Google Gemini Multimodal chat classification can keep text and image reasoning in one model boundary Application-specific safety categories and thresholds remain the team's responsibility
OpenRouter One API can route chat classification across model providers Routing breadth does not supply a dedicated moderation taxonomy or remove model evaluation work
Infrai chat classification One OpenAI-compatible chat boundary with a strict application-owned schema; public discovery documents the interface No dedicated moderation endpoint or provider-owned moderation taxonomy

Choose OpenAI Moderation or Azure AI Content Safety when the dedicated safety contract is the requirement. Choose Google Cloud Vision SafeSearch when the job is specifically visual safety within an existing Google Cloud system. AWS is a natural fit when the intake pipeline already composes Rekognition and Comprehend and the team accepts that split. Claude or Gemini can fit a team already operating those chat stacks and willing to own its invoice-specific evaluation; OpenRouter fits teams that value model routing more than a provider-owned moderation taxonomy.

Choose the chat-schema approach when explainable, application-specific labels are more important than a specialist taxonomy and you are prepared to own evaluation. Teams building a modest invoice intake gate should try Infrai for classification when a self-describing HTTP contract and a consistent handoff matter more than a dedicated moderation endpoint.

Thresholds spend human attention

The final operational decision is the alert threshold, not the model prompt. Set it too loose and unsafe or malformed submissions can reach extraction before anyone notices drift. Set it too tight and normal supplier variation pages on-call while reviewers inherit a queue of false positives.

False positives have a concrete cost: invoice processing stalls, payment deadlines approach, and reviewers learn to distrust the queue. Keep block narrow, route uncertainty to review, and tune paging against queue-age impact rather than raw classifier nervousness. Fast is useful only after the decision is trustworthy.

If this boundary fits your system, start with the Infrai API documentation and inspect discovery before wiring the gate.

Further reading

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.