Dev.to AI 🤖 Ai 👁 0 📖 3 min read

System design for physical AI: live wildfire ignition risk for every distribution feeder, on the utility's own hardware

This is the engineering summary of an open reference architecture. The paper, its object model as JSON and the model register are free to reuse under CC BY 4.0: https://muhammadumar89.github.io/codeninja-research/wildfir

This is the engineering summary of an open reference architecture. The paper, its object model as JSON and the model register are free to reuse under CC BY 4.0: https://muhammadumar89.github.io/codeninja-research/wildfire-risk-distribution-us/ (DOI 10.5281/zenodo.23119325). The operator is an illustrative scenario, not a customer.

A member-owned electric distribution cooperative in the United States runs more than 9,000 miles of overhead line across fire country. Vegetation contact and failing equipment are its leading ignition risks, and they build up between patrol visits. The question it cannot answer today is which feeder segments are most likely to ignite, and which will be exposed to fire weather in the next 48 hours.

The signals exist, spread across systems that never meet: SCADA, GIS, the AMI head end, the outage management system, pole inspection spreadsheets, work management, wildfire cameras, and weather and mesonet feeds. Here is how they turn into a system.

1. Join first, model second

Every source enters through a read-only adapter. The adapters write fourteen typed objects: substation, feeder, feeder segment, pole, recloser, meter, pole inspection record, outage event, feeder segment ignition risk score, red flag warning, wildfire camera station, field crew, work order and public safety power shutoff (PSPS) decision record.

The feeder segment is the focal object. Start from one pole with a defect and one traversal reaches its inspection photos, the segment that carries it, that segment's risk score and the signals behind it, the recloser protecting it and its fast-trip state, the meters and their last gasp history, the red flag warning over the area, the cameras that can see it, and the approved work order with its crew. A document store would need a hand-built join for every hop.

The object model ships as JSON (ontology/objects.json, format hyper-ontology/1) so you can load it instead of redrawing it.

2. Two tiers, placed by the link

Tier Hardware Runs Why there
Edge Fanless, sealed industrial boxes in NEMA 3R/4 enclosures at substations, on patrol trucks and at the yard RF-DETR detection of smoke, downed or leaning poles, vegetation encroachment and hot spots Cellular coverage drops in exactly the storm that matters, so detection cannot depend on the link
Central One ruggedised server of eight 141 GB HBM-class GPUs at the operations center The live distribution model, the risk store, GLM 5.2 under MIT for the work surface The cooperative owns the weights and the record they reason over

3. Size the central node from the weights

weights = 753B parameters x 1 byte (FP8)       = 753 GB
need    = 753 GB x 1.2 (KV cache, activations)  = 904 GB
node    = 8 x 141 GB                            = 1,128 GB -> ~224 GB headroom

The published BF16 weights, about 1.5 TB, would need sixteen cards. FP8 halves that to one node. The time series store beside it is sized for 15-minute AMI intervals from every meter and every recloser, with store-and-forward buffering.

4. Keep a person on every de-energisation

Every PSPS and every fast-trip change is a decision record: a recommendation, a rationale, a named approving operator and the regulator notification reference. It moves through recommended, approved, declined, de-energised and re-energised under that person's hand. The design never opens or closes a recloser, and a work order reaches a crew only after approval in the system crews already use.

5. What it costs

Owning the stack for three years, including an allowance of 63 edge nodes, comes to about 841,000 US dollars with support and power at the Texas industrial price. Renting the same capacity around the clock costs 0.97 to 1.92 million dollars, so ownership is about four fifths of AWS's deepest three-year commitment. A closed frontier model by the token matches the owned stack at about 41 users. Every price is cited in the paper's Appendix A.

Full design, figures and the object model: the paper.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.