Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 15 min read

A CrewAI crew writes our daily pharma brief for three cents

A CrewAI crew writes our daily pharma brief for three cents One daily Pharma brief from a three-agent CrewAI crew costs $0.030 on Claude Haiku 4.5 when the tool hands the model the news API's key sentences: 14,452 prom

A CrewAI crew writes our daily pharma brief for three cents

A CrewAI crew writes our daily pharma brief for three cents

One daily Pharma brief from a three-agent CrewAI crew costs $0.030 on Claude Haiku 4.5 when the tool hands the model the news API's key sentences: 14,452 prompt tokens, 3,042 completion tokens, 4 LLM calls, 8 stories with 8 real links. The same crew on headlines alone costs $0.018 and writes a brief with 1.3 numbers in it; on full article bodies it costs $0.067 and the brief is not more specific (6.3 numbers against 9.0). We ran each configuration three times on the same 40 articles, and this post is the crew, the tool, and the bill.

A CrewAI news agent is a crew of role-playing agents in which one agent owns a tool that fetches articles and the others turn its notes into a document. This post is for people who build agents in Python and want a brief that arrives every morning with a known cost. The industry is Pharma because it is busy (311 English articles tagged Pharma in the 24 hours we used) and because a pharma analyst is a real reader; swap one id and it is any industry. Everything was run on 2026-09-17 with CrewAI 1.15.22; the code, the 12 briefs, the per-call token log and the article snapshot sit next to this post.

Takeaways

  • Key sentences (summary[]) instead of full bodies: 14,452 prompt tokens instead of 53,378, $0.030 instead of $0.067 per brief, and 9.0 specific numbers in the brief instead of 6.3.
  • Headlines only are cheapest at $0.018, but the brief has 1.3 numbers in it and the writer fills the gap with significance it made up.
  • 12 runs, 12 drafts, 0 invented URLs: all 8 links in every draft, 96 in total, were URLs the tool had returned. The editor's job turned out to be cutting, not fact-checking: 181 β†’ 151 words, 24 β†’ 21 links on headlines.
  • The researcher reads 63% of the tokens on headlines, 79% on key sentences and 94% on bodies. The writer and editor never see an article.
  • Share one LLM object between three agents and crew.usage_metrics reports 19,713 prompt tokens for a run that used 6,571. One object per agent.
  • Claude Sonnet 4.6 on the same key sentences: $0.096 a brief, 8.0 numbers. No gain here.

Step 1. Pick the beat and the fields, not the whole feed

The crews in most CrewAI tutorials search the web and hand the model whatever comes back. Ours calls one endpoint with one industry id, so the corpus is fixed and countable. APITube tags every article with industries; GET /v1/suggest/industries?prefix=pharma returns 819 Pharma among others, and the 24 hours from 2026-09-16 06:00 UTC held 311 English articles with that tag.

The other decision is fl=, the field list. An article's body is 4,202 characters at the median and 57,716 at the longest in this day's set; its summary is five key sentences with a sentiment score each, about 781 characters. We ask for both so the tool can hand the model either.

curl -s "https://api.apitube.io/v1/news/everything?industry.id=819&language.code=en&is_duplicate=0&published_at.start=2026-09-16T06:00:00Z&sort.by=published_at&sort.order=desc&per_page=200&fl=id,title,href,source.domain,published_at,summary,body" \
  -H "X-API-Key: $APITUBE_API_KEY"

One result, cut to two of its five key sentences (the first one in this article is a broken line from a financial table, which is what a summary extractor gives you on a press release; the model coped):

{
  "title": "Innate Pharma Reports First Half 2026 Business Update and Financial Results – CIOCoverage- Driven for Technology Leaders",
  "href": "https://www.ciocoverage.com/innate-pharma-reports-first-half-2026-business-update-and-financial-results/",
  "source": {"domain": "ciocoverage.com"},
  "published_at": "2026-09-17T05:22:07.000Z",
  "summary": [
    {"sentence": "β€œ2026 continues to be an important year of execution for Innate, marked by our strategic partnership with Sobi and the strengthening of our financial position,”said Jonathan Dickinson, CEO of Innate Pharma.",
     "sentiment": {"score": 0.13, "polarity": "positive"}},
    {"sentence": "β€œWith the TELLOMAK-3 Phase 3 study initiated, we are targeting the first patient in the study in Q1 2027 as we work toward a filing for accelerated approval in SΓ©zary syndrome.",
     "sentiment": {"score": 0.06, "polarity": "positive"}}
  ],
  "body": "..."
}

Step 2. The tool: one call, three renderings

The tool takes no arguments. It fetches the last 24 hours, keeps the 40 newest articles with at most two per site (stocknews.com alone had 16 that day), and renders them in one of three modes set by NEWS_MODE. Everything else in the crew stays the same, which is what makes the three modes comparable.

# brief_crew.py (excerpt)
import datetime as dt, os, requests
from collections import Counter
from crewai.tools import tool

INDUSTRY_ID = int(os.environ.get("INDUSTRY_ID", "819"))
MAX_ARTICLES = int(os.environ.get("MAX_ARTICLES", "40"))

def fetch_articles():
    now = dt.datetime.now(dt.timezone.utc)
    r = requests.get("https://api.apitube.io/v1/news/everything",
        headers={"X-API-Key": os.environ["APITUBE_API_KEY"]},
        params={"industry.id": INDUSTRY_ID, "language.code": "en", "is_duplicate": 0,
                "published_at.start": (now - dt.timedelta(hours=24)).strftime("%Y-%m-%dT%H:%M:%SZ"),
                "sort.by": "published_at", "sort.order": "desc", "per_page": 200,
                "fl": "id,title,href,source.domain,published_at,summary,body"}, timeout=60)
    r.raise_for_status()
    return r.json()["results"]

def pick(articles, n, per_domain=2):
    seen, out = Counter(), []
    for a in articles:
        d = a["source"]["domain"]
        if seen[d] >= per_domain:
            continue
        seen[d] += 1
        out.append(a)
        if len(out) == n:
            break
    return out

def render(articles, mode):
    lines = []
    for i, a in enumerate(articles, 1):
        lines.append(f"{i}. {a['title']} β€” {a['source']['domain']}, {a['published_at'][11:16]} UTC\n   {a['href']}")
        if mode == "summary":
            lines += [f"   β€’ {s['sentence']}" for s in a.get("summary") or []]
        elif mode == "body":
            lines.append("   " + (a.get("body") or "").replace("\n", " "))
    return "\n".join(lines)

@tool("industry_news")
def industry_news() -> str:
    """Every article about the industry from the last 24 hours: number, title,
    source site, time, URL and, depending on configuration, the key sentences
    or the full text. Call it once; it returns the same set within a run."""
    return render(pick(fetch_articles(), MAX_ARTICLES), os.environ.get("NEWS_MODE", "summary"))

For 40 articles the tool's output is 8,819 characters on headlines, 41,984 on key sentences and 222,931 on bodies. The docstring is the tool description the model reads, so "call it once" is in there, and in all 12 runs it was called once.

Step 3. Three agents, three tasks, one LLM object each

The crew is sequential: the researcher calls the tool and shortlists eight stories with their URLs, the writer turns the shortlist into a brief, the editor removes anything without a URL from the research notes. context=[research] is how a task sees an earlier task's output; the editor gets both the notes and the draft.

from crewai import Agent, Crew, Process, Task

MODEL = os.environ.get("MODEL", "anthropic/claude-haiku-4-5-20251001")

def build_crew(make_llm=lambda: MODEL):
    # One LLM object per agent: crew.usage_metrics sums each agent's counters,
    # so one object shared by three agents reports every call three times.
    researcher = Agent(role="Pharma news researcher",
        goal="Pick the 8 stories a pharma analyst must see today and keep the URL of each",
        backstory="You read the whole feed once and shortlist. You never invent a story or a link.",
        tools=[industry_news], llm=make_llm(), max_iter=5, verbose=False)
    writer = Agent(role="Brief writer",
        goal="Turn the shortlist into a brief a busy reader finishes in two minutes",
        backstory="You write one line per story: what happened, why it matters, then the source link.",
        llm=make_llm(), max_iter=3, verbose=False)
    editor = Agent(role="Brief editor",
        goal="Ship a brief in which every claim is traceable to a URL from the research notes",
        backstory="You cut anything without a source link and anything the notes do not support.",
        llm=make_llm(), max_iter=3, verbose=False)

    research = Task(agent=researcher,
        description="Call industry_news once. From what it returns, shortlist the 8 most consequential "
                    "stories for a pharma analyst (regulatory decisions, trial results, deals, safety, "
                    "pricing). Skip stock-picking pieces and duplicates of the same event.",
        expected_output="A numbered list of 8 items. Each item: headline, one sentence on why it "
                        "matters, and the exact URL from the tool output on its own line.")
    write = Task(agent=writer, context=[research],
        description="Write today's Pharma brief from the research notes. One short paragraph of 1–2 "
                    "sentences per story, ending with the story's URL as a markdown link. No intro, no outro.",
        expected_output="A markdown brief of at most 250 words with one link per story.")
    edit = Task(agent=editor, context=[research, write],
        description="Edit the draft. Remove any story whose URL does not appear in the research notes "
                    "and any fact the notes do not contain. Keep it under 200 words. Keep the markdown links.",
        expected_output="The final brief in markdown, under 200 words, every story ending with a link.")
    return Crew(agents=[researcher, writer, editor], tasks=[research, write, edit],
                process=Process.sequential, verbose=False)

The comment about one LLM object per agent is not style. Crew.calculate_usage_metrics() loops over the agents and adds each agent's llm.get_token_usage_summary() to the total. We ran the crew once with a single object passed to all three agents: it made 4 calls and used 6,571 prompt tokens; crew.usage_metrics reported successful_requests=12 and 19,713 prompt tokens, exactly 3.0Γ—. Passing a model string, or calling a factory as above, gives each agent its own counter. The check is in shared_llm_check.txt.

Step 4. Kick off and print the bill

crew = build_crew()
result = crew.kickoff()
print(result.raw)                       # the brief
u = crew.usage_metrics
cost = u.prompt_tokens / 1e6 * 1.00 + u.completion_tokens / 1e6 * 5.00   # Haiku 4.5 list price
print(f"{u.successful_requests} LLM calls, {u.prompt_tokens:,} in, {u.completion_tokens:,} out, ${cost:.4f}")

result.tasks_output holds the three task outputs in order, so the research notes and the writer's draft are there next to the editor's final. That is where the link check below comes from: every URL in a draft or final is compared with the set of href values the tool returned in that run.

What one brief costs, measured

Twelve kickoffs, three per configuration, all replaying the same articles.json so the only variable is what the tool hands the model. Prices are Anthropic's list prices on 2026-09-17: $1 per million input and $5 per million output tokens for Claude Haiku 4.5, $3 and $15 for Claude Sonnet 4.6. Wall time is 33–43 seconds per brief on Haiku.

Tool hands the model Prompt tokens Completion tokens Cost per brief (min–max) Final brief, words Numbers in the brief Links verified
Headlines only 6,626 2,360 $0.018 ($0.0181–$0.0188) 151 1.3 21 of 21
Key sentences (summary[]) 14,452 3,042 $0.030 ($0.0284–$0.0307) 238 9.0 24 of 24
Full bodies 53,378 2,811 $0.067 ($0.0653–$0.0705) 188 6.3 22 of 22
Key sentences, Claude Sonnet 4.6 14,864 3,458 $0.096 ($0.0935–$0.1000) 236 8.0 24 of 24

Claude Haiku 4.5 unless stated; mean of 3 runs on the same 40 articles. "Numbers in the brief" counts figures outside URLs (counts, percentages, dollar amounts, dates) in the editor's final. "Links verified" sums unique URLs over the 3 finals that match an href the tool returned.

One Pharma brief costs $0.018 on headlines, $0.030 on key sentences, $0.067 on full bodies and $0.096 on key sentences with Sonnet 4.6; min–max across 3 runs is under a cent wide in every configuration

Every run made exactly 4 LLM calls: the researcher twice (once to decide to call the tool, once with the tool's output in front of it), the writer once, the editor once. No retries, no format errors, in 12 of 12. That is why the min–max band is under a cent wide in every configuration: the only thing that moves between runs is how long the model's answers are.

The researcher's share of prompt tokens is 63% on headlines, 79% on key sentences and 94% on full bodies; the writer and editor stay at 918–2,074 tokens each

The researcher is the whole bill. Its second call carries the tool output: 3,700 prompt tokens on headlines, 10,930 on key sentences, 49,615 on bodies. The writer's prompt is 918–1,218 tokens and the editor's 1,537–2,074 in every mode, because they read the researcher's notes, not the articles. If you want a cheaper crew, the lever is what the tool returns, not the number of agents.

What the tool should hand the model

Key sentences put 7.7 more specific numbers in the brief than headlines for 1.1 cents more; full bodies put fewer numbers in the brief than key sentences, 6.3 against 9.0, for 3.8 cents more

The same story in two briefs. From a headlines run: "Genentech's Lunsumio-based regimen achieved its Phase III primary endpoint in follicular lymphoma, advancing a new treatment option that strengthens competitive positioning in oncology." Nothing after the comma is in the headline; the writer was asked for "why it matters" and supplied it. From a key-sentences run: "showing statistically significant reduction in disease progression risk for relapsed/refractory patients, paving the way to convert accelerated approval to full approval and expand second-line indications." Every clause is in the article's five key sentences. Across the three key-sentence runs the briefs carried the $75 million Sobi deal, the Q1 2027 enrollment target, the PACIFIC-9 readout, the 139% rise in ADHD prescriptions and the 55 patients in the JAMA study; we checked each against the article text.

Full bodies did not buy more of that. The body-mode briefs averaged 6.3 numbers to the key-sentence briefs' 9.0, at 2.27Γ— the cost. Three runs is not a large sample and the gap may be noise, but the direction is clear enough for a rule: the right input for a CrewAI news agent is the API's key sentences, not the article bodies, because the key-sentence brief carried 9.0 specific numbers for $0.030 and the body brief 6.3 for $0.067. In order of cost:

  1. Headlines when a title is all the reader needs: a link list, $0.018 a brief, 6,626 prompt tokens.
  2. Key sentences for a brief that states facts: $0.030, 14,452 prompt tokens, 9.0 numbers per brief.
  3. Full bodies only when the writer must quote or the API has no summary, and with a cap on the count: $0.067, 53,378 prompt tokens, 49,615 of them in one call.

Unlike the srk.ai and LangChain tutorials, which hand the model raw search results and never count, this crew pays a predictable 14,452 prompt tokens per brief on key sentences (14,177–14,593 across the three runs), which means a month of daily briefs is $0.89 on key sentences, $0.55 on headlines and $2.02 on bodies.

What the editor actually changes

The editor was written as a fact-checker. In these 12 runs it did not need to be one: all 8 URLs in every draft, 96 across the 12 runs, matched an href the tool had returned, so there was nothing to remove for being invented. What it removed was length. Draft to final: 181 β†’ 151 words and 24 β†’ 21 links on headlines, 267 β†’ 238 words and 24 β†’ 24 links on key sentences, 218 β†’ 188 words and 24 β†’ 22 links on bodies. The 200-word cap it was told to enforce held in 6 of 12 finals; the longest final was 311 words.

It did earn its keep once, in an earlier batch of nine runs we kept in batch-1/. One writer draft carried https://www.fijanzen.ch/... for finanzen.ch, a one-letter slip inside a 148-character URL, and the editor's final had the correct link, restored from the research notes. That is the case for the third agent: not hallucinated stories, but copy errors in long strings, which a diff against the notes catches.

The stop sequence is doing more than you think

CrewAI drives tools for a model without native function calling through text: the agent writes Thought:, Action: industry_news, Action Input: {}, and CrewAI is supposed to stop generation at \nObservation:, run the tool, and paste the result in. It applies \nObservation: as a stop word on the agent's LLM for the duration of the run, and the API cuts generation there.

We ran the model through a backend without server-side stop sequences and cut the text ourselves, so we saw what the model writes when nobody stops it. Every one of the 12 first calls kept going past Action Input: {}. Haiku's ran to 131–475 tokens, of which CrewAI's cut kept 66–75; the rest was mostly a note that it was waiting for the tool. Sonnet 4.6's ran to 5,089–5,925 tokens and up to 19,356 characters: an Observation: containing a made-up feed, with an FDA approval of oral semaglutide and a reuters.com URL that appear nowhere in the day's 311 articles, and a complete final brief built on it. The cut kept 57–63 tokens of those calls, and the real tool output went in after it. If you ever swap in a provider that ignores stop, or you write your own BaseLLM.call(), that invented feed is what reaches your writer.

How we measured, and what is different from your setup

The machine that produced these numbers has a Claude Code subscription and no API key, so cli_llm.py wraps the claude CLI as a CrewAI BaseLLM: same model ids, same prompts, supports_function_calling() returning False, tokens fed to _track_token_usage_internal() so crew.usage_metrics reports them as it would for the API. Three corrections keep the numbers honest: the CLI's fixed preamble (366–370 tokens, measured before each batch and constant across probes) is subtracted from every call; extended thinking is off; and completion tokens for the one over-running call per run are scaled to the share of text CrewAI kept, so the bill is what the API charges with the stop sequence in place. Raw counts and every raw response are in calls.csv and calls/. With LLM(model="anthropic/claude-haiku-4-5-20251001") and an API key you get native tool calling instead of the text protocol, which removes the over-run entirely; the prompt then carries the tool schema and Anthropic's tool-use system prompt, 496 tokens for Haiku 4.5 by its pricing page, instead of CrewAI's text instructions.

Run it on your beat

  1. Look up the industry: GET /v1/suggest/industries?prefix=<word> and put the id in INDUSTRY_ID.
  2. Leave NEWS_MODE=summary and MAX_ARTICLES=40; that is the $0.030 configuration.
  3. Run python brief_crew.py once and read the last line. The prompt-token count follows the tool's output (6,626 tokens on 8,819 characters of headlines here, 53,378 on 222,931 characters of bodies); if it is far above what your mode should give, look at what the tool returned before you tune an agent.

Three agents, one tool, one API call, and a bill of three cents: that is the whole CrewAI news agent, and every number above is in the CSVs next to this post.

FAQ

How do I create a custom tool in CrewAI?

A custom tool in CrewAI is a Python function decorated with @tool("name") from crewai.tools; its docstring becomes the description the agent reads and its signature becomes the input schema. The industry_news tool above takes no arguments and returns a string of 8,819 to 222,931 characters, which CrewAI passes to the model as the observation.

Can CrewAI fetch real-time news?

CrewAI can fetch real-time news through a custom tool that calls a news API; the agent itself has no network access. The tool in this post calls APITube's /v1/news/everything with industry.id=819 and published_at.start set 24 hours back, which returned 311 English Pharma articles for one day, and hands the model the 40 newest with at most two per site.

How much does a CrewAI agent cost to run?

A three-agent CrewAI news crew costs $0.030 per brief on Claude Haiku 4.5 when the tool returns key sentences, $0.018 on headlines and $0.067 on full bodies, measured as crew.usage_metrics times Anthropic's list price over 3 runs each; on Claude Sonnet 4.6 the key-sentence brief costs $0.096. The researcher accounts for 63–94% of the tokens because it is the only agent that reads the articles.

What is the difference between CrewAI and LangChain agents?

CrewAI describes work as agents with roles and tasks with context= dependencies and runs them sequentially or hierarchically, while a LangChain agent is one model in a tool-calling loop and LangGraph makes that loop an explicit graph. For a fixed pipeline like research β†’ write β†’ edit, the CrewAI version is the three Task objects above, and crew.usage_metrics gives the token bill without extra code.

How do I pass tool output between agents in CrewAI?

Tool output moves between CrewAI agents through task context, not through the tool: a task lists earlier tasks in context=[...] and receives their outputs as input. Here the researcher's shortlist of 8 stories with URLs reaches the writer through context=[research]; the writer never calls the tool and never sees the 40 articles, which is why its prompt is 918–1,218 tokens in every mode.

Resources

Disclosure: APITube is our product and the news API used here; the free tier at apitube.io covers a daily brief. The model prices and the CrewAI behaviour are the same whatever feed you plug into the tool.

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.