Dev.to Security ๐Ÿ” Cybersecurity ๐Ÿ‘ 0 ๐Ÿ“– 4 min read

Your cloud security tool is sorting by the wrong number

Every cloud security tool I have used opens on the same screen: a list of findings, sorted by severity, with the criticals at the top in red. It feels like a priority queue. It is not one. Here is the problem with it, u

Every cloud security tool I have used opens on the same screen: a list of findings, sorted by severity, with the criticals at the top in red. It feels like a priority queue. It is not one.

Here is the problem with it, using two findings I had in front of me last month.

Finding A. A critical CVE, CVSS 9.8, remote code execution, on a container image. The image runs in a private subnet. No ingress, no public load balancer, no NAT gateway on the route table. The task role can read one S3 bucket of build artifacts. Nothing else in the account can reach it.

Finding B. A medium misconfiguration, CVSS 5.3. An EC2 instance with IMDSv1 still enabled. The instance sits behind an ALB that is open to the internet and runs an app with a path traversal bug nobody had got round to fixing. Its instance profile can assume a role in the production account, and that role can read the database credentials.

Sorted by severity, A is at the top and B is twelve screens down. Sorted by what an attacker can actually do, B is the only one that matters this week.

Severity is a property of the vulnerability, not of your environment

This is the whole issue in one line. CVSS scores a vulnerability in the
abstract. It has no idea whether the affected host is reachable, what the compromised identity can do next, or whether the blast radius is one build artifact or your customer database.

That is not a criticism of CVSS. It was never meant to be a work queue. The specification says so. We turned it into one because it was the only number every tool agreed on, and because sorting by it is easy to implement.

EPSS helps, because it estimates the probability a vulnerability gets exploited in the wild in the next thirty days. KEV helps more, because it is a list of things that definitely are being exploited right now. Both are still properties of the vulnerability. Neither one knows anything about your network.

What actually makes something urgent

Three things, and none of them are in the CVE record:

Reachability. Can anything an attacker controls get to this asset? That means route tables, security groups, NACLs, load balancers, peering, and whatever your service mesh is doing. In practice most critical findings in a well segmented account are not reachable from anywhere an attacker starts.

Identity. If this asset is compromised, what can the attacker do next? An instance profile that can call sts:AssumeRole into another account is worth more to an attacker than a dozen RCEs on isolated hosts. Over permissioned roles are the real lateral movement path in cloud, not kernel exploits.

Blast radius. What is on the other side of that movement? There is a
difference between reaching a staging bucket and reaching the production
database, and no severity score captures it.

Put those three together and findings stop being a list. They become a graph, and the question changes from "what is the highest score" to "which of these chain together into a route to something that matters".

Doing it yourself

You can get a surprising distance with the APIs you already have. Roughly:

  1. Pull your inventory. ec2:DescribeInstances, ecs:ListTasks, rds:DescribeDBInstances, the equivalents on Azure and GCP.
  2. Pull network configuration. Route tables, security groups, NACLs, load balancer listeners and target groups. Build the reachability edges from that rather than from tags, because tags lie.
  3. Pull IAM. Roles, attached and inline policies, trust policies, instance profiles. The trust policy is the edge you care about, because it is what makes cross account movement possible.
  4. Pull your vulnerability data from wherever it already lives, and join it onto the inventory by instance id or image digest.
  5. Mark your crown jewels by hand. Nothing automatic gets this right, and it is a short list.
  6. Walk the graph from every internet reachable node and see which crown jewels you can reach, and what you had to exploit on the way.

Step 6 is where it gets interesting and where it gets slow. The graph is not large by graph standards, but the path search is not a simple shortest path, because each edge has a precondition: you only traverse the IAM edge if you first got code execution on the node that holds the role.

I built this twice by hand before deciding it was a product. If you want to see what it looks like when it is finished, that is what we do at
Secorvia: a live graph across AWS, Azure, Google Cloud and DigitalOcean, with findings ranked by exploitability and blast radius instead of by CVSS. The scanning is read only, and the free tier does not expire. Build it yourself if you have the time, though. You will understand your own environment far better than any tool will explain it to you.

The thing I got wrong for a long time

I used to think the goal was to get the critical count to zero. It is not, and chasing it actively hurts, because you spend your week on unreachable criticals and the reachable mediums stay open.

A better target: zero findings that sit on a path from the internet to anything you would be sad to lose. That number is usually small enough to actually fix, which is the entire point.

If you do this differently, I would genuinely like to hear it, particularly how you handle reachability through service meshes. That is the part I am least happy with.

๐Ÿ“ฐ Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes โ€” full credit and traffic to the original publisher.