Dev.to AI 🤖 Ai 👁 0 📖 4 min read

Building Effective Agentic Workload Systems

Building an agentic workload system means wiring an LLM to tools and letting it iterate until it reaches a decision. In this tutorial we will ship an autonomous incident response agent that reads an alert, queries mock t

Building an agentic workload system means wiring an LLM to tools and letting it iterate until it reaches a decision. In this tutorial we will ship an autonomous incident response agent that reads an alert, queries mock telemetry, and emits a concrete operations recommendation. The pattern works for any domain where you need structured reasoning over unreliable data.

What you'll need

Step 1: Define the tools and Oxlo.ai client

I define three simulated functions that stand in for real infrastructure APIs. Each returns JSON so the model can parse results reliably. I initialize the client against Oxlo.ai because its flat per-request pricing keeps costs predictable even when I pass large telemetry blobs back into the conversation on every turn.

import json
from openai import OpenAI

client = OpenAI(base_url="https://api.oxlo.ai/v1", api_key="YOUR_OXLO_API_KEY")

def get_metrics(service: str, duration_min: int = 5):
    return json.dumps({
        "cpu_percent": 96.4,
        "memory_percent": 88.2,
        "error_rate": 14.3,
        "p99_latency_ms": 2450
    })

def get_recent_deploys(service: str, limit: int = 3):
    return json.dumps([
        {"sha": "a1b2c3d", "time": "10:04 UTC", "message": "update payment gateway timeout"}
    ][:limit])

def run_diagnostics(service: str, check_type: str):
    if check_type == "db_connection_pool":
        return json.dumps({"status": "degraded", "open_connections": 198, "max_connections": 200})
    return json.dumps({"status": "unknown"})

TOOL_MAP = {
    "get_metrics": get_metrics,
    "get_recent_deploys": get_recent_deploys,
    "run_diagnostics": run_diagnostics,
}

TOOLS = [
    {
        "type": "function",
        "function": {
            "name": "get_metrics",
            "description": "Fetch recent metrics for a service.",
            "parameters": {
                "type": "object",
                "properties": {
                    "service": {"type": "string"},
                    "duration_min": {"type": "integer", "default": 5}
                },
                "required": ["service"]
            }
        }
    },
    {
        "type": "function",
        "function": {
            "name": "get_recent_deploys",
            "description": "List recent deployments for a service.",
            "parameters": {
                "type": "object",
                "properties": {
                    "service": {"type": "string"},
                    "limit": {"type": "integer", "default": 3}
                },
                "required": ["service"]
            }
        }
    },
    {
        "type": "function",
        "function": {
            "name": "run_diagnostics",
            "description": "Run a specific diagnostic check.",
            "parameters": {
                "type": "object",
                "properties": {
                    "service": {"type": "string"},
                    "check_type": {"type": "string", "enum": ["db_connection_pool", "disk_space", "network_latency"]}
                },
                "required": ["service", "check_type"]
            }
        }
    }
]

Step 2: Write the system prompt

The system prompt constrains the model to an ops persona and forces a machine-readable final answer. Keeping instructions explicit reduces hallucinated actions.

SYSTEM_PROMPT = (
    "You are an autonomous incident response agent. "
    "When you receive an alert, investigate by calling tools. "
    "After gathering evidence, output a final message with exactly this format:\n\n"
    "ACTION: \n"
    "REASON: \n\n"
    "Do not emit the final action until you have called at least one tool. "
    "Prefer ROLLBACK if a recent deploy coincides with elevated error rates. "
    "Prefer SCALE_UP if metrics show sustained high CPU or memory without a clear code change cause. "
    "Otherwise, PAGE_ONCALL."
)

Step 3: Build the agent loop

Now the loop. I send the alert to the model with the tool definitions. If the model requests tool calls, I execute them, append the results, and prompt again. I cap iterations to prevent runaway traces. I use qwen-3-32b here because Oxlo.ai highlights it for agent workflows, and the flat per-request cost means extra context from tool results does not inflate the bill on every turn.

def run_agent(alert_text: str, service: str = "payment-api", max_steps: int = 5):
    messages = [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": f"Service: {service}\nAlert: {alert_text}"}
    ]

    for _ in range(max_steps):
        response = client.chat.completions.create(
            model="qwen-3-32b",
            messages=messages,
            tools=TOOLS,
            tool_choice="auto",
            temperature=0.2,
        )

        message = response.choices[0].message

        assistant_msg = {
            "role": "assistant",
            "content": message.content or "",
        }
        if message.tool_calls:
            assistant_msg["tool_calls"] = [tc.model_dump() for tc in message.tool_calls]
        messages.append(assistant_msg)

        if not message.tool_calls:
            return message.content

        for tc in message.tool_calls:
            fn_name = tc.function.name
            fn_args = json.loads(tc.function.arguments)
            result = TOOL_MAP[fn_name](**fn_args)

            messages.append({
                "role": "tool",
                "tool_call_id": tc.id,
                "name": fn_name,
                "content": result,
            })

    return "Reached max steps without final action."

Step 4: Run it

Here is the entry point. I feed the agent a realistic latency spike alert and print the final recommendation.

if __name__ == "__main__":
    alert = "payment-api p99 latency spiked to 2.4s, error rate above 10%"
    print(run_agent(alert))

When I run this, the agent first calls get_metrics and get_recent_deploys, sees the correlation between the gateway timeout deploy and the elevated error rate, and returns:

ACTION: ROLLBACK
REASON: Error rate and latency spiked immediately after the payment gateway timeout deploy at 10:04 UTC.

Next steps

Swap the mock functions for real Datadog or Kubernetes API calls, and add a Slack webhook so the agent posts directly to your incidents channel. Because Oxlo.ai uses flat per-request pricing, stuffing large log responses back into the context window does not inflate your bill on every turn, which makes long agent traces predictable. You can compare plans at https://oxlo.ai/pricing.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.