ForgeZero and Gloria: Build Tooling for People Who Still Want to Know What Their CPU Is Doing
I'm Alex Voste, the person behind ForgeZero and its resident experimental language, Gloria. I'm not going to open with a mission statement. I'm going to open with the actual reason this exists: I got tired of trusting to
I'm Alex Voste, the person behind ForgeZero and its resident experimental language, Gloria. I'm not going to open with a mission statement. I'm going to open with the actual reason this exists: I got tired of trusting tools that refuse to tell me what they're doing.
Modern tooling has a bad habit of treating "it worked" as a substitute for "I understand why it worked." Garbage collectors show up in your hot path uninvited. Build systems recompile the world because a timestamp looked at them funny. Compilers hide behind Makefiles that hide behind generators that hide behind more generators, until nobody on the team can actually explain what make is going to do before it does it. None of this is malicious. Most of it is genuinely convenient for most people. It is also, for anyone doing real systems work, a tax — paid in CPU cycles, paid in debugging hours, paid in that specific flavor of dread when a build fails and the error message is a lie by omission.
So I built two things instead of complaining about one.
ForgeZero: A Build System That Refuses to Pretend
In 1976 a Bell Labs engineer named Stuart Feldman got annoyed at a coworker who kept shipping broken binaries — editing a source file and forgetting to rebuild whatever depended on it. His fix became make, a tool built to answer one question correctly, every time: what actually needs rebuilding right now, and what doesn't?
Fifty years later we're still answering that question badly, and most engineers have simply stopped noticing. Make's timestamp-and-phony-target model leaked complexity into anything bigger than a weekend project. Autotools buried that complexity under a generation-of-generators architecture that remains one of the most feared corners of open source. CMake showed up to generate other build systems instead of solving the underlying problem — which is either a clever abstraction or a quiet confession, depending on which decade of CMake broke your afternoon. Ninja, to its credit, went the other direction entirely: radically minimal, explicitly not meant to be hand-written, a fast and dumb execution target for something smarter sitting above it.
None of these were ever built for the workload I actually live in — assembly-heavy code, hundreds of small C translation units, embedded and bare-metal targets, mixed-language repos where coordinating the compilers costs more than running them. ForgeZero exists to close that gap. The thesis is one sentence:
A build system should handle planning, invalidation, and routing. It should not try to reinvent the compiler.
ForgeZero is not GCC, Clang, Zig, NASM, FASM, GAS, or LD, and has zero interest in becoming any of them. It's the thin, transparent coordination layer deciding what those tools need to do, when, and whether they need to do it again at all — then getting out of the way. The entire value proposition compresses into one relationship:
Build Throughput ≈ Useful Compiler Work
─────────────────────────────────────────
Discovery + Dispatch + Repeated Work + Invalidation Cost
Every optimization in ForgeZero attacks the denominator. The numerator — actual code generation — stays permanently someone else's job, because that's a division of labor worth respecting.
Cache identity, done the way people usually get wrong
The classic homegrown-build-cache mistake is treating "the source bytes didn't change" as sufficient proof that a compiled object is still valid. It isn't. The same source file produces a different correct binary under a different compiler, a different version, a different target triple, a different sysroot, different flags, different sanitizer policy. ForgeZero's cache identity is built from sorted inputs, their digests, the action name, and the full toolchain context — not the source hash alone.
Metadata-first invalidation checks file size and mtime before touching a single byte of content. Only a mismatch triggers a full keyed digest. The on-disk format, FZHC3, is written temp-file-then-rename, specifically so a crash never leaves you a half-written cache entry quietly lying to you later.
Two hashes, one boundary nobody's allowed to blur
BoomBoom Hash is the incremental-invalidation engine, built over keyed BLAKE3. Build context — compiler identity, version, target triple, flags — gets domain-separated into a derived key, so identical source bytes under different build contexts never collide. Source files split into logical chunks sitting as leaves in a flat Merkle tree; change one chunk and only its path to the root recomputes — O(log n), not a full re-aggregation.
BB64 exists because BoomBoom Hash, built on a cryptographic primitive, was doing more work than some invalidation paths actually need. It's a purpose-built, non-cryptographic 64-bit hash with a hand-written AVX2 hot path and a scalar fallback verified bit-exact against it. On a Ryzen 7 PRO 4750U it measured roughly 16.87 GB/s at zero allocations per op — about 40x FNV-1a and nearly 10x BLAKE3, for cache-key generation that never needed cryptographic guarantees in the first place.
And here's the part worth actually trusting: BB64 is never used anywhere security-sensitive. Integrity checks, package verification, cryptographic surfaces — all of that stays on keyed BLAKE3, permanently. Collapsing that boundary to save one dependency is exactly the kind of shortcut that turns into a CVE two years later. We didn't take it.
Below the Go standard library
Source discovery on Linux skips filepath.WalkDir for raw getdents64 via unix.Getdents, paired with openat and O_NOFOLLOW so traversal can't silently follow a symlink somewhere it shouldn't. Metadata goes through a single statx call, falling back to Fstatat. Directory reads retry on EINTR. Path ordering is deterministic, which matters enormously the first time you're chasing a flaky build that "worked yesterday." Manual dirent parsing cut Linux discovery allocations on a 1,024-file benchmark from roughly 3,182 to 2,088 per operation — a third fewer, on the code path that runs first, every single build.
io_uring integration batches file reads into one submission cycle instead of one syscall per file, drains completion queues with explicit bounds checking, and handles zero-length files without generating a garbage READ request. And then the detail that actually tells you how the project thinks: IORING_OP_GETDENTS was never wired up, because the target kernel ABI doesn't expose a usable opcode for it — so directory listing stayed on the syscall path where it actually works, instead of faking a feature-completeness checkbox. Unglamorous. Also exactly the decision that separates infrastructure you can trust from infrastructure that looks impressive in a README until someone runs it on the wrong kernel.
On the linker side, worker and scheduler queue counters sit on separate 64-byte cache lines on purpose, to kill false sharing — the scenario where two unrelated atomics land on the same cache line and start fighting over coherency traffic under load, tanking throughput for reasons invisible anywhere in the code. That's the kind of optimization that only shows up in codebases actually profiled under real concurrency, not just "compiles and passes CI."
The Redis number, read honestly
fz -p performance -dir redis/src -mode c -out redis-server -j 16
user 27.66 s
system 11.68 s
CPU 1155%
elapsed 3.404 s
1155% CPU utilization means the work genuinely spread across hardware threads, not hidden behind a fast single-thread path. 3.4 seconds elapsed against nearly 40 seconds of combined CPU time puts the bottleneck in parallel compile-and-link work, which is exactly where you want it. The dependency closure wasn't trivial either — fpconv, hdr_histogram, hiredis, libxxhash, linenoise, lua, and tre each got compiled into static archives and handed to the linker alongside Redis's own units. This is end-to-end orchestration latency on one specific host, toolchain, and cache state — not a claim about compiler speed, and it means nothing without the hardware, governor, and cache-warmth that produced it. We say that ourselves, unprompted, because a number without context is just marketing with extra steps.
Gloria: The Language That Refuses to Hide Anything
Here's the part I actually want systems people to sit with.
Gloria isn't trying to compete with C, Rust, or a full assembler frontend. It occupies a narrow, deliberate niche: a minimal language for producing a small, fully-understood fragment of machine code in contexts where a libc, a runtime, and a general-purpose compiler pipeline are simply overkill. Gloria compiles straight to native AMD64 binaries a few hundred bytes in size, with no dependency on libc or any managed runtime. There's no garbage collector, no comfortable abstraction layer deciding what "actually happens" on your behalf — because every layer between your intent and the instruction stream is a layer that can lie to you about what's real.
Zero-allocation by default. No heap sitting behind the scenes being polite about it. Gloria talks to the kernel directly through syscalls — mmap, sys_open, write — which means memory control starts on line one, not somewhere after three layers of runtime initialization you didn't ask for.
Syntax built around how the hardware actually behaves, not around how a 1970s compiler wanted you to phrase things: >> chaining with register passing, |? guard operators for error propagation without a pyramid of nested checks, repeat loops that read the way you'd say them out loud, native vectors backed by REP MOVSQ, hash tables on open addressing.
Coroutines in pure assembly. Cooperative green threads (fibers) with register-level context switching and 4KB mmap-backed stacks, natively. That's not a small feature to get right, and most "systems-adjacent" languages don't bother attempting it at all.
The pipeline is embarrassingly honest about itself:
source -> lexer -> tokens -> codegen -> machine bytes -> relocation -> raw binary
No mandatory object file. No libc startup sequence. No standard linker script. The emitter writes bytes directly into a contiguous output buffer; calls get written first with relocation placeholders, and once the function table is complete, relative displacements get patched straight into the buffer:
disp32 = offset(target) - (offset(call) + 5)
That's the entire mechanism letting Gloria skip a full object-file linker for small raw images. Compile this:
fn main() {
let a = 10;
let b = 20;
if b > a {
print("hello world");
}
}
and disassembling the resulting binary shows exactly what you'd hope for from a language that doesn't hide things: an entry jump and stack-frame prologue, 64-bit immediate loads for the literals, a cmp/jle pair lowering the comparison — and, because the observed output path writes directly to 0xb8000 (the classic x86 VGA text-memory address), a byte-by-byte character-and-attribute write loop instead of a hidden printf call buried three abstractions deep. Every opcode in that disassembly traces back to a specific, readable decision the emitter made. That's the whole point of Gloria existing. Not competing with production languages. Refusing to lie to you about what your code turned into.
Right now, filesystem interaction is solid — reading and writing data confidently, not a toy demo. The active front line is dynamic arrays (vectors), which isn't a feature for its own sake — it's the foundation everything downstream, language and compiler both, actually needs.
Next up is full socket support, which cracks the door open to servers, network tooling, and game engines written in Gloria with nothing hiding between your code and the packets actually hitting the wire. And the real north star — the thing all of this is actually pointed at — is self-hosting. The day ForgeZero's Gloria compiler is itself written in Gloria is the day the whole project proves its own premise. A language still borrows credibility from whatever built it until it can bootstrap itself.
I Use AI. It Doesn't Get a Vote.
I keep a neural net around as a sparring partner — kicking around syntax ideas, stress-testing a design before committing to it, occasionally helping me phrase a post like this one without sounding like I've been awake for thirty hours (I frequently have been). What it doesn't get is unsupervised trust. My relationship with AI-generated code is closer to "overconfident intern I review like a hostile auditor" than "pair programmer I defer to." It accelerates thinking. It doesn't do the thinking, and it goes nowhere near the parts of this project where a mistake means a corrupted binary instead of an awkward sentence.
What We Deliberately Didn't Build
The most senior-engineer-coded part of this whole project isn't a feature list — it's the short list of things we explicitly chose not to ship:
- BB64 never touches a security-sensitive path. Integrity and verification stay on keyed BLAKE3, permanently.
-
IORING_OP_GETDENTSwas never faked into existence to check a feature-completeness box the kernel ABI doesn't actually support. -
clone3/execveatweren't forced into the process launcher without a real lifecycle design behind them. - A shared-memory persistent cache wasn't shipped without a versioned format, crash recovery, and an ownership protocol — a cache that can silently corrupt itself is worse than no cache.
- No thousand-x speedup claims without workload-specific numbers standing behind them.
Any project can list what it built. Fewer are willing to publish, in plain language, what they refused to ship half-finished. That restraint is a stronger signal than another benchmark chart.
Come Argue With Us
If you genuinely enjoy getting into the hardware, obsessing over a single instruction more than is strictly healthy, and want to watch something real and unapologetically low-level get built in public instead of reading the polished retrospective six months later — this is the invitation.
GitHub: github.com/forgezero-cli
Live updates, real bugs, actual wins, Telegram: t.me/forgezeroteam
Direct contact: [email protected]
No garbage collectors were harmed in the making of this project. Several were, however, deeply and personally offended.
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.