Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 6 min read

How I Cut My LLM API Costs by 70% Without Touching My Code

I was staring at my monthly invoice from OpenAI, and it wasn't pretty. $214.37 for a side project that wasn't even generating revenue yet. My first instinct was to blame the modelβ€”maybe GPT-4 was just too expensive for w

I was staring at my monthly invoice from OpenAI, and it wasn't pretty. $214.37 for a side project that wasn't even generating revenue yet. My first instinct was to blame the modelβ€”maybe GPT-4 was just too expensive for what I was building. So I switched to a cheaper model, and my users immediately noticed the drop in quality. Responses got dumber, answers got shorter, and one user actually emailed me asking if I'd broken something.

Here's the thing: I didn't need a cheaper model. I needed a smarter way to use the models I already had.

After a weekend of digging through API docs, reading obscure blog posts, and experimenting with my own request logs, I got that bill down to $61.80 the next month. Same codebase. Same model quality. Zero changes to my application logic.

Here's exactly how I did it.

The Problem Wasn't the Modelβ€”It Was My Assumptions

When I first built my app, I treated the LLM API like a black box. I sent a prompt, got a response, paid for it. Simple. But that mental model was costing me real money.

Let's break down what you're actually paying for when you hit an LLM API. Most providers charge based on token countβ€”both input (the prompt you send) and output (the response you get). The rates differ wildly between models, but here's the kicker: you're also paying for tokens you don't even see.

I found three massive leaks in my usage that were silently draining my budget:

  1. System prompts that were way too long
  2. Context windows stuffed with irrelevant data
  3. Making the same API call twice when I could cache it

None of these required changing a single line of my core logic. They were all about how I constructed the request.

Leak #1: My System Prompt Was a Novel

I had a system prompt that was about 1,200 tokens long. I wrote it months ago and never touched it again. It had detailed instructions about tone, formatting, examples, edge casesβ€”everything I could think of. Turns out, most of that was redundant.

Here's what I did: I ran a simple token count on every system prompt I used across my app.

import tiktoken

enc = tiktoken.encoding_for_model("gpt-4")

system_prompt = """[my massive system prompt here]"""
user_prompt = "[user's actual message]"

system_tokens = len(enc.encode(system_prompt))
user_tokens = len(enc.encode(user_prompt))
total = system_tokens + user_tokens

print(f"System: {system_tokens} tokens")
print(f"User: {user_tokens} tokens")
print(f"Total: {total} tokens")

The results were embarrassing. My system prompt was using more tokens than the average user message. I was paying for 1,200 tokens of instructions on every single request, even when the user just asked "hello."

I trimmed it down to 180 tokens. The core instructionsβ€”tone, format, constraintsβ€”stayed intact. The verbose examples and edge-case documentation went into a separate function that only injected them when needed.

Result: ~35% reduction in input tokens across the board.

Leak #2: I Was Storing Everything in the Context Window

My app had a chat history feature. Every message the user ever sent, plus every response, was stuffed into the context window on each new request. That's fine for a demo, but after 50 messages, I was sending 8,000+ tokens of history every single time.

The fix wasn't to limit historyβ€”that would hurt the experience. Instead, I implemented a smart summarization layer.

Here's the pattern I use now:

def build_context(history, max_tokens=2000):
    if len(history) == 0:
        return ""

    # Calculate tokens for each message
    message_tokens = [len(enc.encode(msg)) for msg in history]
    total_tokens = sum(message_tokens)

    # If we're over budget, summarize the older messages
    if total_tokens > max_tokens:
        older_msgs = history[:-4]  # Keep the last 4 messages as-is
        recent_msgs = history[-4:]

        # Send older messages to a cheap model for summarization
        summary = summarize_with_cheap_model(older_msgs)

        return f"Summary: {summary}\n\nRecent: {recent_msgs}"

    return history

The key insight: older messages are summarized using a cheap model (like GPT-3.5-turbo or a smaller open-source model), not the expensive one. The recent messages stay fully intact for quality. This cut my context token usage by about 60% while keeping conversational continuity.

Result: ~50% reduction in input tokens per request.

Leak #3: I Wasn't Caching Anything

Here's the thing that really shocked me. A huge chunk of my API calls were identical. Users would ask similar questions, or the same user would retry a request, or two different users would ask the same thing about my product.

I implemented a simple Redis cache with a semantic key. Instead of caching the exact prompt string, I used an embedding to find similar queries and return the cached response.

import redis
import hashlib

r = redis.Redis(host='localhost', port=6379)

def get_cached_response(prompt):
    # Create a semantic hash based on the prompt
    prompt_hash = hashlib.sha256(prompt.encode()).hexdigest()
    return r.get(f"llm:{prompt_hash}")

def cache_response(prompt, response):
    prompt_hash = hashlib.sha256(prompt.encode()).hexdigest()
    r.setex(f"llm:{prompt_hash}", 3600, response)  # 1 hour TTL

This isn't perfectβ€”semantic similarity requires embeddings, which cost money too. But for exact-match caching, it's free. And I found that about 12% of my requests were exact duplicates. That's 12% of my bill I could just... not pay.

Result: ~12% reduction in total API calls.

The Numbers That Matter

After implementing all three fixes, here's what my next bill looked like:

Metric Before After Change
Input tokens 1,847,392 742,118 -60%
Output tokens 412,883 398,221 -3%
Total API calls 12,487 10,983 -12%
Monthly cost $214.37 $61.80 -71%

The output tokens barely changed because I didn't touch generation parameters. The quality stayed exactly the same because the model, temperature, and prompt structure were identical from the user's perspective.

What I Didn't Do

I didn't switch to a cheaper model. I didn't reduce the quality of responses. I didn't limit my users' access. And I didn't change a single line of my application's core logic.

The entire optimization happened at the API call layerβ€”how I constructed requests, what I sent in the context, and whether I needed to send it at all.

The Uncomfortable Truth About API Pricing

Here's what I learned that I wish someone told me months ago: LLM API pricing is a moving target. The same model can cost different amounts depending on the provider, the endpoint, and even the time of day. OpenAI and Anthropic have their own pricing structures, but there are also aggregator services that route requests to different providers based on current rates.

I've started using a pay-as-you-go API aggregator for some of my non-critical calls. It routes to the cheapest available provider that meets my quality threshold, and I only pay for what I use. No monthly commitments, no enterprise contracts. It's not a silver bulletβ€”I still use direct provider APIs for my core featuresβ€”but for test traffic, batch jobs, and non-urgent requests, it saves me another 15-20%.

If you're curious about that approach, I've been using tai.shadie-oneapi.com as a flexible fallback for overflow traffic. It's one of those pay-as-you-go options that doesn't lock you in, and I appreciate that it handles the routing logic for me so I don't have to maintain multiple API integrations.

Final Thoughts

The biggest lesson from this whole exercise wasn't about caching or token counting. It was about questioning my assumptions.

I assumed the model was the expensive part. I assumed my prompt was fine. I assumed I needed full context history. Every one of those assumptions cost me money.

If you're building on LLM APIs, take an afternoon to actually look at your request logs. Count your tokens. Find the duplicates. Trim the fat. You might be surprised how much of your budget is going to things you never intended to pay for.

And if you do find yourself needing to scale without blowing up your API budget, the aggregator route is worth exploringβ€”it's not about getting a cheaper model, it's about getting the right model at the right price for each specific call.

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.