System design for physical AI: predicting container terminal truck turn time with edge vision and an open-weight model
This is the engineering summary of an open reference architecture. The paper, its object model as JSON and the model register are free to reuse under CC BY 4.0: https://muhammadumar89.github.io/codeninja-research/truck-t
This is the engineering summary of an open reference architecture. The paper, its object model as JSON and the model register are free to reuse under CC BY 4.0: https://muhammadumar89.github.io/codeninja-research/truck-turn-container-terminal-us/. The operator is an illustrative scenario, not a customer.
A container terminal in the United States runs two berths, eight ship to shore cranes, twenty six rubber tyred gantry cranes (RTGs) over fourteen yard blocks, eleven truck gate lanes and an on dock rail ramp. Truck turn time averages 54 minutes and spikes above 90 on export peaks. Nobody can say why while it is happening, because the causes live in different systems: the terminal operating system, the gate system, two crane vendors' telemetry, reefer monitoring, rail switch lists, cameras, weather and tide.
Here is how that turns into a system.
1. Join first, model second
Every source enters through one of three adapter families (integration, EDI, and an edge node for cameras), never directly. The adapters write twelve typed objects: vessel call, container, yard block, RTG, ship to shore crane, truck visit, gate lane, rail cut, reefer plug, yard person, transfer zone and safety event.
The links carry verbs: a yard block assigns an RTG, a truck visit enters through a gate lane, a vessel call discharges to containers, a reefer plug powers a container. Start from one truck visit that ran long and one traversal reaches the gate exception that held it, the container's yard block, the RTG assigned there and its fault codes, the vessel call and its discharge order, and the rail cut the box may miss. A document store holds every record and answers none of that.
The object model ships as JSON (ontology/objects.json, format hyper-ontology/1) so you can load it instead of redrawing it.
2. Three tiers, three clocks
| Tier | Hardware | Runs | Its clock |
|---|---|---|---|
| Edge | Fanless IP-rated enclosures at the yard blocks, gate and quay | RF-DETR detection, a Roboflow tracker on CPU | The camera frame: the transfer-zone conflict verdict never waits on a network |
| Site | One GPU node, L40S-class 48 GB or H100-class 80 GB | Chronos-2 turn time forecast, Qwen3-Embedding-0.6B retrieval | The operation: recomputed as crane cycles, gate reads and appointments arrive |
| Frontier | One node, eight 141 GB HBM cards | GLM 5.3 at FP8 for the planner work surface, served by vLLM or SGLang | The human: interactive sessions in front, batch reconciliation behind |
3. Size each tier from the weights
edge RF-DETR Nano to Large ~61-68 MB at 16 bit -> several streams per accelerator
site Chronos-2 ~0.24 GB (16 bit) + embedder ~0.6 GB (8 bit) -> one card, room for 32K activations
frontier 753B x 1 byte (FP8) = 753 GB; x 1.2 = 904 GB; 8 x 141 GB = 1,128 GB -> one node, ~10 kW
Pin the edge serving runtime (ONNX Runtime, OpenVINO or Triton) after a bench measurement on the yard's own streams. A detector sized from a datasheet is the first way a vision system disappoints.
4. Keep the boundary closed
The model holds no outbound connection. Detection lives at the edge because no safety reflex should cross a network hop. Footage of longshore labour never leaves the site; the only things that cross tiers are detections, forecasts and the records a person acts on.
5. Keep a person on every action
Surfaces warn and propose. The named planner acts inside the terminal operating system, and the named safety supervisor acknowledges every safety event, with the disposition written to the record. Write-back into the terminal operating system is a second phase, gated on labour and IT sign-off.
6. What it costs
Owning the stack for three years comes to about 722,000 US dollars with support and power at the US industrial price. Renting the same GPUs around the clock costs 1.12 to 2.52 million dollars, so ownership is about two thirds of the cheapest three-year cloud commitment. A closed frontier model by the token matches the owned stack at about 35 users and costs more for every user after that. Every price is cited in the paper's Appendix A.
Full design, figures and the object model: the paper.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.