Dev.to WebDev 🛠 Dev 👁 0 📖 6 min read

How to Build Your Own Search Aggregator in 100 Lines of Code

How to Build Your Own Search Aggregator in 100 Lines of Code Still bouncing between arxiv, crossref, and pubmed for academic research or tech scouting? Copy-pasting the same keyword three times? This article shows you

How to Build Your Own Search Aggregator in 100 Lines of Code

Still bouncing between arxiv, crossref, and pubmed for academic research or tech scouting? Copy-pasting the same keyword three times? This article shows you how to build your own search aggregator with the SeekAll SDK in about 100 lines — you write the rules, the data runs locally, and the server touches nothing.

The Pain: Information Scattered Across Silos

When I do academic research or tech scouting, I keep running into the same scene:

  • Searching a paper keyword means checking arxiv for preprints, crossref for DOI metadata, and pubmed for biomedical literature
  • Searching a tech stack means checking GitHub for repos, Hacker News for discussion, and Stack Overflow for Q&A
  • The same keyword, copied across 3–5 sites, with results scattered across different tabs

The "xxx aggregator search" sites out there are either wrappers around search engines, stuffed with ads, or limited to a single data source. More importantly — your search terms pass through someone else's server, making privacy and compliance a black box.

What I wanted was something that is:

  1. Rule-customizable: I decide which sources to search by writing my own rules
  2. Local-first: Search terms never pass through any intermediary server
  3. Tool-neutral: The platform ships zero default data sources, avoiding compliance risk

So I built SeekAll — a rule-engine SDK + a rule marketplace + BaaS, with 0 rules by default, all rules running on the user's own machine.

Up and Running in 3 Lines

npm i @seekall/sdk @seekall/rule-arxiv @seekall/rule-crossref
import { createEngine } from "@seekall/sdk";
import arxiv from "@seekall/rule-arxiv";
import crossref from "@seekall/rule-crossref";

const engine = createEngine({ rules: [arxiv, crossref] });
const hits = await engine.search("transformer attention mechanism");
hits.forEach((h) => console.log(`${h.title}\n  ${h.url}\n`));

That's it. createEngine takes an array of rules, engine.search queries all of them concurrently, and returns a unified Hit[] structure. Your search terms only ever run inside the Node process on your machine — SeekAll's server has zero contact with your search content.

Hands-On: Writing a GitHub Trending Rule (10 Lines)

The core of SeekAll is "rules". A rule is simply an object implementing the Rule interface, whose key method is search: it takes a keyword + context and returns Hit[].

Let's write a 10-line rule using GitHub Trending as an example:

import type { Rule, Hit, RuleContext } from "@seekall/sdk";

export const githubTrendingRule: Rule = {
  name: "@my-org/rule-github-trending",
  version: "0.1.0",
  riskLevel: "L1", // L0 academic-pure / L1 general open-source / L2 community-reviewed / L3-L4 high-risk

  async search(keyword: string, ctx: RuleContext): Promise<Hit[]> {
    const url = `https://api.github.com/search/repositories?q=${encodeURIComponent(keyword)}&sort=stars&order=desc`;
    const r = await ctx.fetch(url, {
      headers: {
        Accept: "application/vnd.github.v3+json",
        "User-Agent": "SeekAll/0.5",
      },
    });
    const data = await r.json();
    return (data.items || []).slice(0, 10).map((item: any) => ({
      title: item.full_name,
      url: item.html_url,
      snippet: item.description,
      meta: { stars: item.stargazers_count, language: item.language },
    }));
  },
};

That's a complete rule. The key points:

  1. name: the rule's unique identifier, in npm package-name format
  2. riskLevel: the risk rating. SeekAll uses a 5-level scale (L0–L4): L0 is academic-pure sources like arxiv/crossref, L1 is general open-source like GitHub, and L3/L4 are high-risk sources (admin-only, never public)
  3. search: an async function that takes keyword + context and returns Hit[]
  4. ctx.fetch: rules use the SDK-provided ctx.fetch instead of bare fetch, so the SDK can uniformly handle concurrency control, timeouts, and caching

Why ctx.fetch instead of fetch? Because SeekAll's performance is tier-based:

Tier Concurrency Timeout Cache
free 3 10s none
trial (¥1/7 days) 5 8s none
monthly (¥18/30 days) 10 5s 5min
lifetime (¥68/100 yrs) 20 3s 5min

ctx.fetch automatically rate-limits + caches based on your license tier — the rule code itself doesn't need to care about any of this.

Publish to npm + Submit to the SeekAll Marketplace

Once your rule is written, two steps make it usable by others:

1. Publish to npm

# in the rule package directory
pnpm build
npm publish --access public

A rule is just a regular npm package — anyone can install it with npm i @your-org/rule-xxx.

2. Submit to the SeekAll Rule Marketplace

SeekAll has a rule marketplace (https://seekall.winmelon.cn/rules) where users can browse + subscribe to rules. The submission flow:

  1. Register an account on SeekAll
  2. Click "Submit Rule" in the marketplace
  3. Fill in the rule name, npm package name, risk level, and description
  4. Wait for community review (L0–L2) or admin final review (L3–L4)
  5. Once approved, the rule appears in the marketplace list and other users can subscribe

Important: SeekAll's marketplace does not host rule code — it only does listing + subscription. Rule code lives on npm, and users pull it directly from npm when installing. This keeps SeekAll's server at zero contact with resource content, preserving tool neutrality.

5-Level Risk Rating: The Tool-Neutrality Philosophy

The most central design in SeekAll is the 5-level risk rating:

Level Description Visibility Examples
L0 Academic-pure Public arxiv, crossref, pubmed
L1 General OSS Public GitHub, Hacker News, Stack Overflow
L2 Community-reviewed Public (after review) specific community APIs
L3 High-risk Admin-only gray-area resource sites
L4 Extreme-risk Admin-only never public

Why this design? Because the tool itself is neutral, but data sources are not. A knife can chop vegetables or hurt people — SeekAll chooses to be the "knife", not the "knife shop": we provide the engine, you decide what to search.

  • L0–L2 rules are public; anyone can install them
  • L3–L4 rules are never visible to non-admins, even if you pay
  • The server never calls resource sites (there's no axios/fetch/http in apps/api/src/modules/rule/); all requests originate from the user's machine

This means SeekAll, as a platform, touches no resource content. The compliance boundary is clear.

Performance Differentiation: Free Is Enough, Pay to Go Faster

SeekAll is a commercially-minded project, not purely open source. The SDK core is AGPL-3.0 open source, but performance differentiation requires a license:

  • free: 3 concurrency + 10s timeout, for light use
  • trial ¥1/7 days: 5 concurrency + 8s timeout, experience the full feature set
  • monthly ¥18/30 days: 10 concurrency + 5s timeout + 5min cache, for heavy use
  • lifetime ¥68/100 years: 20 concurrency + 3s timeout + 5min cache, one-time buyout

Licenses are sold through the WM card shop (card code + webhook activation) — no payment SDK is integrated.

Why not fully free? Because running a rule marketplace (review, takedown, DMCA handling) costs money. A paywall also filters out some abuse.

Summary: 100 Lines to Your Own Search Aggregator

Full code recap:

// 1. import the SDK + rules
import { createEngine } from "@seekall/sdk";
import arxiv from "@seekall/rule-arxiv";
import crossref from "@seekall/rule-crossref";
import { githubTrendingRule } from "./my-rules/github-trending"; // your own rule

// 2. create the engine
const engine = createEngine({
  rules: [arxiv, crossref, githubTrendingRule],
});

// 3. search
const hits = await engine.search("transformer attention mechanism");

// 4. handle results
hits.forEach((h) => {
  console.log(`[${h.source}] ${h.title}`);
  console.log(`  ${h.url}`);
  console.log(`  ${h.snippet?.slice(0, 100)}`);
  console.log();
});

Counting the GitHub Trending rule you wrote yourself (10 lines), that's roughly 30 lines of core code. Add error handling, deduplication, and output formatting, and 100 lines is plenty to build your own search aggregator.

How this differs from "xxx aggregator search" sites:

Aggregator search site SeekAll
Search terms via The site's server Only on your machine
Data sources Built-in, you can't choose You write the rules, search what you want
Compliance Borne by the site Tool-neutral, platform touches nothing
Extensibility Wait for the site to update Publish npm packages, community-contributed

CTA

If you also do academic research or tech scouting and are tired of bouncing between sites, and want to build a search aggregator that's truly yours, SeekAll is currently the most neutral option. You write the rules, the data runs locally, and the server touches nothing.

This is the first article in the SeekAll tutorial series. The next one, "Why I Don't Build a Search Website, Only an SDK", will discuss the tool-neutrality design philosophy.

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.