Dev.to WebDev πŸ›  Dev πŸ‘ 0 πŸ“– 16 min read

Building PrismJudge: An Offline-First, Mathematically Rigorous Hackathon Platform

An Engineering Deep Dive into Building PrismJudge Targeting the Grand Prize ($800 USD) and Best Judging Engine ($100 USD) Repository: https://github.com/jhanikhilnath/PrismJudge_DogFood_Hack Stack: Node.js 22 LTS Β· TypeS

Building PrismJudge: An Offline-First, Mathematically Rigorous Hackathon Platform

An Engineering Deep Dive into Building PrismJudge

Targeting the Grand Prize ($800 USD) and Best Judging Engine ($100 USD)

Repository: https://github.com/jhanikhilnath/PrismJudge_DogFood_Hack

Stack: Node.js 22 LTS Β· TypeScript Β· Fastify v5 Β· Embedded SQLite 3 (WAL Mode) Β· Vanilla Studio Light SSR

Official Acceptance Status: 7 / 7 PASS in run.py Β· 81 / 81 PASS in npm test Β· 641 / 641 PASS in Route Audit Β· 89 / 89 PASS in Container E2E

Table of Contents

  • 1. Executive Summary & The Core Philosophy
    • The One-Command Rule
    • Comprehensive Verification Scorecard
  • 2. System Architecture: The In-Process Fastify & SQLite Pipeline
    • In-Process Request Flow
    • Why Embedded node:sqlite Beats PostgreSQL for Hackathon Portals
  • 3. Mathematical Rigor: The Judging & Normalization Engine
    • 3.1 Empirical Bayesian Shrinkage Z-Score Normalization
    • 3.2 Formal Proof: The Singularity Resolution of Judge jdg_07
    • 3.3 Bradley-Terry Pairwise Model via Minorization-Maximization (MM)
    • 3.4 Inter-Rater Reliability: Intraclass Correlation Coefficient ICC(1,1)
  • 4. Security Architecture & Adversarial Defenses
    • 4.1 The Hard Peer Isolation Defense
    • 4.2 Temporal Submission Deadline Guard
    • 4.3 Anti-Sybil Community Voting Guard
  • 5. Engineering Challenges & Battle Stories: What Broke and How We Fixed It
    • Challenge 1: The Browser Caching Deadlock & Persona Switcher Trapping
    • Challenge 2: The Two-Page Diploma Print Overflow
    • Challenge 3: Chronological Skew in the Append-Only Audit Ledger
  • 6. Frontend Design
  • 7. Verification & Testing Runbook
    • The Official Acceptance Report
  • 8. Conclusion & Lessons Learned

1. Executive Summary & The Core Philosophy

The prompt for DOGFOOD 2026 was deceptively simple, yet brutally unforgiving:

"Build the platform that will judge you."

In 72 hours, teams were challenged to construct an entire end-to-end hackathon management and evaluation portal capable of ingesting messy legacy fixtures (41 projects, 30 judges, 8 tracks, and 126 reviews), enforcing strict submission deadlines, isolating evaluator peer scores under adversarial probing, evening out juror bias with documented statistical normalization, and serving traffic completely offline.

Most teams immediately reached for the standard modern web toolbox: Next.js 15, Django or FastAPI, PostgreSQL 16, Redis, and multi-container Docker Compose networks. While functional, that architecture carries immense cognitive and operational overhead: slow cold-boot times (15–30 seconds), container networking race conditions, connection pool exhaustion, and fragile multi-service health checks.

We made an audacious architectural bet: The One-Command Rule.

The One-Command Rule

docker compose up

With the host machine's physical network disconnected, our entire platform boots from a cold start in under 600 milliseconds, mounts an embedded SQLite 3 database in Write-Ahead Logging (WAL) mode, runs idempotent schema migrations, ingests all fixture submissions and scores, seeds four authenticated evaluation personas, and serves production-ready HTTP/1.1 traffic on port 8080.

PrismJudge Editorial Welcome Portal rendering real-time macro telemetry, active competition tracks, and offline platform statusFigure 1: The PrismJudge Welcome Portal, rendering real-time macro telemetry, competitive tracks, and offline platform status.

Comprehensive Verification Scorecard

Competition Dimension Requirement / Invariant Verified Platform Status
Claimed Tiers claimed = ["T1", "T2"] PASS (7/7 in official run.py)
T1 Core Platform Public gallery, fixture rendering, deadline rejection 100% PASS (22 / 22 assertions)
T2 Judging Engine Workload queue, 4-slider rubric, IDOR peer isolation, CSV export 100% PASS (27 / 27 assertions)
T3 Public Balloting Fisher-Yates hash shuffle, anti-Sybil rate limiter, comments stream 100% PASS (7 / 7 assertions)
T4 Stretch Features Bradley-Terry MM arena, OpenAPI 3.1, private verifiable diplomas 100% PASS (33 / 33 assertions)
Internal Unit Suite Automated regression & boundary tests (npm test) 81 / 81 PASS across 12 test suites
Role & Route Crawler Comprehensive HTTP status matrix (tests/audit_script.mjs) 641 / 641 assertions PASS
Live Container Audit Full browser & API container audit (tests/comprehensive_e2e_audit.mjs) 89 / 89 assertions PASS
Diploma Print Layout Single-page landscape PDF export (@media print) Strictly 1 Page (Pages: 1, 0 overflow)
Container Cold Boot Offline Docker initialization time < 600ms on http://localhost:8080

2. System Architecture: The In-Process Fastify & SQLite Pipeline

Rather than splitting our system across microservices, we engineered PrismJudge as a high-throughput, unified monolith inside Node.js 22 LTS using Fastify v5. Fastify was selected over Express and NestJS due to its low overhead (up to 4x faster than Express), built-in JSON schema compilation, and clean plugin lifecycle.

In-Process Request Flow

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                             INCOMING HTTP REQUEST                           β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                       β”‚
                                       β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                      Fastify Global Pipeline Hooks                          β”‚
β”‚  β”œβ”€ onRequest: Fastify Cookie parsing & session secret verification         β”‚
β”‚  β”œβ”€ preHandler: resolveUserHook (maps session token to req.user)            β”‚
β”‚  └─ onSend: Enterprise Security Headers & Cache-Control Enforcement         β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                       β”‚
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β–Ό                              β–Ό                              β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Public Views β”‚              β”‚ Evaluator Queue β”‚            β”‚ Organizer Admin β”‚
β”‚  - Gallery   β”‚              β”‚  - Rubric Model β”‚            β”‚  - Live Console β”‚
β”‚  - Detail    β”‚              β”‚  - Peer Guard   β”‚            β”‚  - Bayesian Proofβ”‚
β”‚  - Ballot    β”‚              β”‚  - Pairwise MM  β”‚            β”‚  - CSV Export   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜              β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜            β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
        β”‚                              β”‚                              β”‚
        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                       β”‚
                                       β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                     Centralized Data Access Layer                           β”‚
β”‚                       (src/db/queries.ts)                                   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                       β”‚
                                       β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚               Node.js 22 Embedded SQLite 3 (node:sqlite)                    β”‚
β”‚      - DatabaseSync singleton with PRAGMA journal_mode = WAL                β”‚
β”‚      - PRAGMA synchronous = NORMAL Β· PRAGMA foreign_keys = ON               β”‚
β”‚      - In-process zero-latency query execution (< 150 microseconds)         β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Why Embedded node:sqlite Beats PostgreSQL for Hackathon Portals

Node.js 22 introduced native, zero-dependency SQLite support via node:sqlite (DatabaseSync). Running SQLite in-process yields decisive advantages for hackathon evaluation:

  1. Zero Connection Latency: Database queries are synchronous C++ function calls directly against the SQLite engine in the same memory address space. Average query execution is under 150 microseconds.
  2. Write-Ahead Logging (WAL Mode): Readers never block writers, and writers never block readers. Concurrency bottlenecks disappear even under intensive test crawls.
  3. Atomic Multi-Table Transactions: Workload redistribution and scoring updates execute inside isolated ACID transactions with zero risk of partial commits.
  4. Offline Resilience: The database is an immutable, single-file artifact (portal.sqlite). There is no Postgres daemon to wait for, no TCP socket handshakes, and no password authentication failures during container boot.
// src/db/index.ts β€” The In-Process SQLite WAL Singleton
import { DatabaseSync } from 'node:sqlite';
import { config } from '../config.js';

let dbInstance: DatabaseSync | null = null;

export function getDatabase(): DatabaseSync {
  if (!dbInstance) {
    dbInstance = new DatabaseSync(config.dbPath);
    // Enable Write-Ahead Logging & Foreign Key integrity
    dbInstance.exec('PRAGMA journal_mode = WAL;');
    dbInstance.exec('PRAGMA foreign_keys = ON;');
    dbInstance.exec('PRAGMA synchronous = NORMAL;');
    dbInstance.exec('PRAGMA busy_timeout = 5000;');
  }
  return dbInstance;
}
PrismJudge Public Submissions Gallery featuring responsive track filter pills, instant client-side search, and project spotlightingFigure 2: The Public Submissions Gallery featuring responsive track filter pills, instant client-side search, and project spotlighting.

3. Mathematical Rigor: The Judging & Normalization Engine

In hackathons, the most common scoring methodology is naive arithmetic averaging: summing reviewer numbers and dividing by the review count. This approach is fundamentally broken:

  • Evaluator Severity Bias (Hawks vs. Doves): A score of 3.5 from a strict judge who rarely awards above 4.0 represents excellence, whereas a 3.5 from a generous judge who routinely gives 5.0 represents mediocrity.
  • Sparse Bipartite Evaluation Graphs: Because no evaluator can review all 41 submissions, different projects encounter different subsets of judges.
  • The Zero-Variance Singularity: What happens when an evaluator awards identical scores to every submission?

To address these statistical realities, PrismJudge implements a three-pillar mathematical scoring engine: Empirical Bayesian Shrinkage Z-Score Normalization, Bradley-Terry Pairwise Minorization-Maximization (MM), and Two-Way Random Effects Inter-Rater Reliability (IRR).

       RAW JUROR SCORES (S_ij)
                 β”‚
                 β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚     Global Prior Computation    β”‚ ───►  ΞΌ_0 β‰ˆ 3.567,  Οƒ_0^2 β‰ˆ 0.2606
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                 β”‚
                 β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚   Empirical Bayesian Shrinkage  β”‚ ───►  Shrinks juror mean ΞΌ_j* and variance (Οƒ_j*)^2
β”‚      (m = 3.0 pseudo-reviews)   β”‚       Guarantees: Οƒ_j* > 0  (Zero-Variance Immune)
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                 β”‚
                 β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚    Standardized Z-Score Scale   β”‚ ───►  Z_ij = (S_ij - ΞΌ_j*) / Οƒ_j*
β”‚  Score_norm = clamp(70 + 12Β·Z)  β”‚       Linear transform into calibrated [0, 100]
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                 β”‚
                 β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚   Bradley-Terry MM Solver       β”‚ ───►  Iterative latent capability Ο€_i
β”‚     (Pairwise Arena Ranks)      β”‚       Elo Rating: R_i = 1500 + 400Β·log10(Ο€_i)
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                 β”‚
                 β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Composite Leaderboard & Export β”‚ ───►  80% Normalized + 20% Pairwise Elo
β”‚       (RFC 4180 CSV Export)     β”‚       Strict descending rank order
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

3.1 Empirical Bayesian Shrinkage Z-Score Normalization

Standard Z-score standardization shifts a judge's scores by their sample mean SΛ‰j​ and scales by their sample standard deviation sj​ :

Zijnaive​=sj​Sijβ€‹βˆ’SΛ‰j​​

When a judge evaluates only 2 or 3 projects, sample variance sj2​ is an extremely noisy estimator. Worse, if a judge gives the same score to all projects (zero variance, sj​=0 ), Zijnaive​ divides by zero and produces NaN or Infinity.

We resolve this by treating the hackathon's global population parameters as an Empirical Bayesian Prior:

  1. Global Prior Estimation:

    ΞΌ0​=N1​i,jβˆ‘β€‹Sijβ€‹β‰ˆ3.567,Οƒ02​=Nβˆ’11​i,jβˆ‘β€‹(Sijβ€‹βˆ’ΞΌ0​)2β‰ˆ0.2606βŸΉΟƒ0β€‹β‰ˆ0.5105
  2. Shrunk Juror Mean ( ΞΌj⋆​ ):

    ΞΌj⋆​=nj​+mnj​SΛ‰j​+mΞΌ0​​
    where m=3.0 represents the prior strength (pseudo-observations). A juror with few reviews is shrunk strongly toward the global mean ΞΌ0​ ; as njβ€‹β†’βˆž , ΞΌj⋆​→SΛ‰j​ .
  3. Shrunk Juror Variance ( (Οƒj⋆​)2 ):

    (Οƒj⋆​)2=max(1,njβ€‹βˆ’1)+mmax(0,njβ€‹βˆ’1)vj​+mΟƒ02​​,Οƒj⋆​=(Οƒj⋆​)2​
    where vj​ is the juror's sample variance.
  4. Normalized Score Transformation:

    Zij​=Οƒj⋆​Sijβ€‹βˆ’ΞΌj⋆​​,Scoreijnorm=clamp(70+12β‹…Zij,0,100)

3.2 Formal Proof: The Singularity Resolution of Judge jdg_07

In fixtures.json, judge jdg_07 evaluated three submissions (prj_13, prj_25, prj_37), assigning identical rubric scores to all three:

S13,7​=3.667,S25,7​=3.667,S37,7​=3.667⟹v7​=0.000

Under a naive Z-score calculation:

Zi,7naive​=03.667βˆ’3.667​→00​=NaN

Under our Empirical Bayesian Shrinkage model:

(Οƒ7⋆​)2=(3βˆ’1)+3.0(3βˆ’1)β‹…0+3.0β‹…(0.2606)​=5.00+0.7818​=0.15636
Οƒ7⋆​=0.15636β€‹β‰ˆ0.3954>0

Because m=3.0>0 and Οƒ02β€‹β‰ˆ0.2606>0 , the numerator is strictly positive:

(Οƒj⋆​)2β‰₯nj​+mβˆ’1mΟƒ02​​>0βˆ€nj​β‰₯1

Conclusion: Division by zero is mathematically impossible. Juror jdg_07 receives a well-conditioned standard deviation Οƒ7β‹†β€‹β‰ˆ0.3954 , allowing normalized scores to be computed smoothly without special-casing or crashing.

Lead Organizer Console displaying macro system telemetry, the Bayesian Shrinkage singularity handled badge for jdg_07, juror calibration metrics, and live standingsFigure 3: The Lead Organizer Console displaying macro system telemetry, the Bayesian Shrinkage singularity handled badge for jdg_07, juror calibration severity metrics, and live competition standings.

3.3 Bradley-Terry Pairwise Model via Minorization-Maximization (MM)

For head-to-head showdown comparisons in the Pairwise Arena, we model submission win probabilities using the Bradley-Terry model. The probability that project i beats project j is parameterized by latent positive capabilities Ο€i​,Ο€j​ :

P(i≻j)=Ο€i​+Ο€j​πi​​

We solve for Ο€ using Hunter's (2004) Minorization-Maximization (MM) algorithm with an uninformative Dirichlet prior ( Ξ±=0.1 ) to ensure connectivity over sparse comparison graphs:

Ο€i(t+1)​=βˆ‘jξ€ =i​πi(t)​+Ο€j(t)​Nij​​+Ξ±MWi​+α​

We then convert latent capabilities into an intuitive Elo rating:

Ri​=1500+400β‹…log10​πi​

Head-to-head Pairwise Arena with dual project cards, keyboard shortcuts for fast voting, and automated round generation


Figure 4: The Head-to-Head Pairwise Arena with dual project cards, keyboard shortcuts (1, 2, T), and automated round generation.

3.4 Inter-Rater Reliability: Intraclass Correlation Coefficient ICC(1,1)

To evaluate consensus across multi-reviewed submissions, we compute the One-Way Random Effects Intraclass Correlation Coefficient ICC(1,1) :

ICC(1,1)=MSB+(kβˆ’1)MSWMSBβˆ’MSW​

where MSB is mean square between submissions, MSW is mean square within submissions, and k is the harmonic mean of reviews per project. Our mathematical audit revealed a fascinating insight: uncalibrated raw scores yield an ICC(1,1) of -0.009 (near zero) because evaluator severity bias confounds within-project residual variance. After Bayesian shrinkage normalization, ICC(1,1) climbs to +0.0304, verifying that our normalization isolates genuine project quality from evaluator noise.

4. Security Architecture & Adversarial Defenses

A portal cannot be production-ready if its access control is purely decorative. Many hackathon implementations hide sensitive buttons in HTML templates while leaving underlying REST APIs wide open.

       CLIENT REQUEST
             β”‚
             β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚     resolveUserHook     β”‚ ───► Extract token from Cookie / Bearer
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
             β”‚
             β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β”‚                                                 β”‚
             β–Ό                                                 β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚    Public Endpoints     β”‚                       β”‚    Protected Endpoints  β”‚
β”‚  - GET /projects        β”‚                       β”‚  - POST /projects/new   β”‚
β”‚  - GET /projects/:id    β”‚                       β”‚  - GET /judge/*         β”‚
β”‚  - GET /vote            β”‚                       β”‚  - GET /organizer/*     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                       β”‚  - GET /api/judge/scoresβ”‚
                                                  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                                               β”‚
                                                               β–Ό
                                                  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                                                  β”‚       RBAC Checks       β”‚
                                                  β”‚  - requireRole(...)     β”‚
                                                  β”‚  - Peer Isolation Check β”‚
                                                  β”‚  - Deadline Guard       β”‚
                                                  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                                               β”‚
                                       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                                       β”‚                                               β”‚
                                       β–Ό (Denied)                                      β–Ό (Allowed)
                        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                        β”‚   HTTP 401 Unauthorized     β”‚                 β”‚    Execute Controller       β”‚
                        β”‚   HTTP 403 Access Restrictedβ”‚                 β”‚   Log to Audit Ledger       β”‚
                        β”‚   (Zero Information Leak)   β”‚                 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

4.1 The Hard Peer Isolation Defense

The official acceptance checker (run.py) specifically tests for Insecure Direct Object References (IDOR):

# run.py Check 5
GET /api/judge/scores?judge=judge_a
Sent with Cookie: session=jdg_b_44de (Judge B)
Expected: HTTP 401 or 403

In PrismJudge, peer isolation is enforced in the preHandler hook before any business logic executes:

// src/core/rbac.ts β€” Hard Peer Isolation Middleware
export async function enforceJudgePeerIsolation(req: FastifyRequest, reply: FastifyReply): Promise<void> {
  if (!req.user) {
    return reply.code(401).send({ error: 'Unauthorized: Authentication session required' });
  }

  // Participants have zero access to judge scoring records
  if (req.user.role === 'participant') {
    return reply.code(403).send({ error: 'Forbidden: Participants cannot inspect evaluation scores' });
  }

  // Organizers have universal administrative audit privileges
  if (req.user.role === 'organizer' || req.user.role === 'admin') {
    return;
  }

  // Evaluators are strictly quarantined to their own score ledger
  if (req.user.role === 'judge') {
    const query = req.query as { judge?: string; judge_id?: string };
    const requestedJudge = query.judge || query.judge_id;

    if (requestedJudge && !isJudgeSelf(requestedJudge, req.user.userId)) {
      logAuditEvent({
        actorId: req.user.userId,
        actorRole: req.user.role,
        action: 'SECURITY_VIOLATION_PEER_PROBE',
        resourceType: 'scores',
        payload: { targetJudge: requestedJudge },
        ipAddress: req.ip,
      });

      return reply.code(403).send({
        error: 'Forbidden: Judges are strictly prohibited from inspecting peer scores',
      });
    }
  }
}

Evaluator Workspace displaying assigned submissions, weighted rubric sliders (40/30/20/10%), and active peer isolation notices


Figure 5: The Evaluator Workspace displaying assigned submissions, weighted rubric sliders (40/30/20/10%), and active peer isolation notices.

4.2 Temporal Submission Deadline Guard

The competition specification requires that late submissions be refused:

// src/routes/projects.ts β€” Server-Side Deadline Enforcement
if (isSubmissionsClosed()) {
  return reply.code(403).send({
    error: 'Submissions closed',
    message: 'The submission deadline has elapsed. No new projects may be submitted.',
    deadline: event.submissions_close,
  });
}

4.3 Anti-Sybil Community Voting Guard

Community voting is susceptible to automated script stuffing. We implemented a multi-layered defense:

  1. Sliding-Window IP Rate Limiting: A rolling 3.0-second delay per IP address. Flooding triggers HTTP 429 Too Many Requests.
  2. Deterministic Ballot Shuffling: Project display orders are randomized using a Fisher-Yates hash shuffle keyed to the voter's session hash. This prevents "positional bias" where the first project receives an unfair advantage.
  3. Database-Level Uniqueness Constraint: community_votes enforces UNIQUE(voter_hash, project_id). Duplicate ballot submissions trigger an immediate HTTP 409 Conflict.

Public Community Ballot featuring randomized ordering, live badge tallies, and duplicate voting prevention


Figure 6: The Public Community Ballot featuring randomized ordering, live badge tallies, and duplicate voting prevention.

5. Engineering Challenges & Battle Stories: What Broke and How We Fixed It

The true measure of an engineering project is not just what went right, but the obscure, hair-pulling bugs encountered along the way and the rigorous engineering required to resolve them.

Challenge 1: The Browser Caching Deadlock & Persona Switcher Trapping

The Symptom

During interactive verification, a frustrating bug surfaced:

"The switching personality via the menu is broken. I can only switch into judge 2 and nothing else. Switching via the route of logout and then log back in works though."

Clicking "Judge B (Wei)" successfully switched the user into Judge 2. But once logged in as Judge 2, clicking "Lead Organizer", "Judge A", or "Participant" in the dropdown menu seemingly did nothingβ€”the page refreshed, and the user remained stuck as Judge 2. Logging out and logging back in via /login worked perfectly.

Authenticated profile dropdown menu where the caching deadlock occurred


Figure 7: The authenticated profile dropdown menu where the caching deadlock occurred.

The Deep-Dive Investigation

We traced the complete request lifecycle across the browser and server:

  1. In a previous iteration, an interceptor had been added to src/public/js/app.js to preserve the user's current URL upon switching personas:
   // Legacy app.js interceptor
   document.querySelectorAll('a[href^="/api/auth/switch/"]').forEach((link) => {
     link.addEventListener('click', () => {
       const curPath = window.location.pathname;
       link.setAttribute('href', `${link.getAttribute('href')}?redirect=${encodeURIComponent(curPath)}`);
     });
   });
  1. When the user navigated to /judge/dashboard as Judge 2, clicking "Lead Organizer" in the menu triggered this script, converting the target link to: /api/auth/switch/organizer?redirect=%2Fjudge%2Fdashboard
  2. On the backend, resolveSmartRedirect() contained this logic:
   // Legacy auth.ts logic
   if (path.startsWith('/judge')) {
     if (targetRole === 'judge' || targetRole === 'organizer' || targetRole === 'admin') {
       return candidate; // candidate = '/judge/dashboard'
     }
     return getRoleDefaultPath(targetRole);
   }
  1. Because organizers have administrative access to judge dashboards, the server concluded that /judge/dashboard was a valid destination for an organizer! It redirected the user right back to /judge/dashboard.
  2. Crucially, Fastify served static assets with default caching (Cache-Control: public, max-age=0), and layout.ejs linked <script src="/static/js/app.js"></script> without a cache-buster. Even after we removed the client-side interceptor in Git, the browser continued executing the stale, cached script from its disk cache!

The Permanent Solution

We applied a three-part hardening fix:

  1. Asset Version Query Busting: In src/views/layout.ejs, we appended ?v=20260930 to both styles.css and app.js. Any change instantly forces browsers to download fresh assets.
  2. Static Asset Revalidation: In src/app.ts, we configured Cache-Control: no-cache, must-revalidate for /static/ assets so browsers always validate ETags.
  3. Cross-Role Redirect Guard: In src/routes/auth.ts, we updated resolveSmartRedirect(): switching to organizer from /judge/* explicitly overrides the candidate and routes to /organizer/dashboard.
// src/routes/auth.ts β€” Hardened Redirection Guard
if (path.startsWith('/judge')) {
  // Only judges remain on /judge/*; organizers route to their console
  if (targetRole === 'judge') {
    return candidate;
  }
  return getRoleDefaultPath(targetRole);
}

We verified this by adding automated unit test 8b to tests/persona_switch.test.ts.

Challenge 2: The Two-Page Diploma Print Overflow

The Symptom

PrismJudge issues private, printable landscape diplomas for team participants and evaluators, complete with gold foil seal medallions, SVG guillochΓ© borders, credential verification numbers, and dual committee signatures.

However, during print testing via Chrome's Print to PDF (Cmd+P), the certificate looked gorgeous on screen but spanned across 2 pages, pushing the cryptographic audit strip onto page 2.

Private Honors Participation Diploma rendered with ornamental framing, gold medallion seal, and credential verification


Figure 8: The Private Honors Participation Diploma rendered with ornamental framing, gold medallion seal, and credential verification.

The Investigation

Using headless Chrome DevTools Protocol (CDP) to generate PDFs and inspecting them with pdfinfo:

  • Page count: Pages: 2
  • Root cause: <main class="app-main"> had a default responsive padding of padding: 2.25rem 1.5rem 4rem; (nearly 100px total). Inside @media print, this padding was not stripped, and large margins on .diploma-wrapper caused vertical content to exceed the physical 210mm height of A4/Letter paper in landscape orientation.

The Permanent Solution

We completely redesigned the @media print CSS block in src/public/css/styles.css:

  1. Defined strict page dimensions: @page { size: landscape; margin: 6mm 8mm; }.
  2. Stripped all outer paddings: html, body, main.app-main { padding: 0 !important; margin: 0 !important; }.
  3. Applied page-break-inside: avoid; break-inside: avoid; across all diploma containers.
  4. Scaled font sizes, signature lines, and the gold medallion to fit within a ~480px total height budget.
/* src/public/css/styles.css β€” Print-Perfect Landscape Stylesheet */
@media print {
  @page {
    size: landscape;
    margin: 6mm 8mm;
  }
  html, body, main.app-main {
    padding: 0 !important;
    margin: 0 !important;
    background: transparent !important;
  }
  .app-header, .app-footer, .cert-actions-bar, #toast-container {
    display: none !important;
  }
  .diploma-wrapper, .diploma-sheet, .diploma-frame-outer, .diploma-frame-inner {
    page-break-inside: avoid !important;
    break-inside: avoid !important;
  }
}

Verification: Headless Chrome PDF generation confirmed:

pdfinfo /tmp/cert.pdf | grep "Pages:"
# Output: Pages: 1

Challenge 3: Chronological Skew in the Append-Only Audit Ledger

The Symptom

During rapid execution of the 81-test regression suite, queries to the audit ledger (GET /api/organizer/audit) occasionally returned log entries in non-chronological order.

The Investigation

Audit records were originally sorted by ORDER BY created_at DESC. In automated testing environments, multiple audit events (login, score submission, assignment modification) occur within the same millisecond. Because ISO 8601 string timestamps lacked microsecond precision, SQLite broke timestamp ties arbitrarily based on table scan order.

The Solution

We updated src/db/queries.ts to order audit logs by ORDER BY rowid DESC. SQLite's 64-bit integer rowid is strictly monotonic and guaranteed to reflect physical insertion order, completely eliminating timestamp collision artifacts.

6. Frontend Design

Many AI-assisted projects suffer from "slop": generic dark-purple neon gradients, unreadable low-contrast gray text, raw JSON token dumps in the UI, and placeholder copy.

PrismJudge enforces strict Design Principles:

  1. Studio Light Typography:
    • Primary: Plus Jakarta Sans with tight negative tracking (-0.03em) on titles.
    • Monospace: JetBrains Mono for all identifiers, scores, dates, and tabular figures.
    • Backgrounds: Clean, warm slate (#F8FAFC, #FFFFFF) with crisp 1px borders (#E2E8F0).
  2. Zero Technical Debris:
    • Raw session tokens (org_7f2a, usr_part_...) are never dumped into UI headers.
    • Debug token boxes and "Copy Token" buttons are excluded from user-facing screens.
  3. Adaptive Local Time Formatting:
    • Timestamps are rendered server-side as standard ISO UTC: <time class="local-time" datetime="2026-03-01T18:00:00Z">Sunday, March 1, 2026 at 6:00 PM UTC</time>
    • A lightweight client script progressively enhances the element using the user's browser timezone (Intl.DateTimeFormat), displaying both local and UTC times.
  4. Privacy-Conscious Roster Displays:
    • Registered team members' email addresses are masked on public project details (ada@***.org) to protect participant privacy while retaining verification clarity.

Project Details View displaying technical breakdown, masked team roster, repository links, and rate-limited discussion comments


Figure 9: Project Details View displaying technical breakdown, masked team roster, repository links, and rate-limited discussion comments.

Official Results Portal with podium winner showcases, track category champions, and toggleable embargo status


Figure 10: The Official Results Portal with podium winner showcases, track category champions, and toggleable embargo status.

Event Settings Console with interactive date-time pickers, rubric weighting, and double-confirmation protection


Figure 11: The Event Settings Console with interactive date-time pickers, rubric weighting, and double-confirmation protection.

Team Management console displaying cryptographic credentials, member rosters, and team invite codes


Figure 12: The Team Management console displaying cryptographic credentials, member rosters, and team invite codes.

7. Verification & Testing Runbook

To guarantee total reliability, PrismJudge is validated through five independent verification layers before every deployment.

# Gate 1: TypeScript typechecking and bundle distribution build
npm run build

# Gate 2: Internal unit and integration regression test suite
npm test
# Result: 81 / 81 PASS across 12 test suites (0 failures, 0 skipped)

# Gate 3: Comprehensive 641-point role and route matrix crawler
node tests/audit_script.mjs
# Result: 641 / 641 assertions PASS across all 4 personas

# Gate 4: End-to-end containerized browser & API verification suite
node tests/comprehensive_e2e_audit.mjs
# Result: 89 / 89 checks PASS against live Docker container

# Gate 5: Official competition acceptance checker
docker compose build && docker compose up -d
python3 run.py .dogfood.toml

The Official Acceptance Report (acceptance-report.txt)

DOGFOOD 2026 acceptance report
portal: http://localhost:8080
claimed: T1 T2
fixtures: fixtures.json

T1  gallery is public ................. PASS
T1  project from fixtures shown ....... PASS
T1  closed event refuses submissions .. PASS
T2  judge sees own scores ............. PASS
T2  judge cannot see peer scores ...... PASS
T2  participant blocked ............... PASS
T2  csv export works .................. PASS

claimed T1 T2, verified T1 T2

8. Conclusion & Lessons Learned

Building PrismJudge for DOGFOOD 2026 reinforced three foundational software engineering truths:

  1. Boring Architecture Wins: In a 72-hour competition, eliminating distributed systems complexity (separate frontend containers, message brokers, external databases) enabled us to move faster, debug deeper, and deliver microsecond query performance. An in-process SQLite WAL database inside Node.js 22 was faster and more reliable than any multi-container setup.
  2. Mathematical Defense is Non-Negotiable: A hackathon platform cannot simply "average the numbers." Implementing Empirical Bayesian Shrinkage solved real-world edge cases like judge jdg_07's zero-variance singularity and transformed evaluator noise into trustworthy rankings.
  3. Adversarial Verification Uncovers the Truth: Hiding buttons in UI templates does not stop an attacker with curl. Writing custom penetration scripts that aggressively probe IDOR endpoints, simulate stale browser caches, and verify print layouts is what separates amateur prototypes from production software.

PrismJudge proves that a modern web application can be offline-first, mathematically sound, beautifully designed, and rock-solid under adversarial testingβ€”all booting in under one second.

Authored for the DOGFOOD 2026 Hackathon Platform Competition.

Repository: https://github.com/jhanikhilnath/PrismJudge_DogFood_Hack

πŸ“° Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.