Fingerprint or Not: Choosing the Right ZoomEye Query for a GeoServer Investigation
Fingerprint or Not: Choosing the Right ZoomEye Query for a GeoServer Investigation Two people can search an internet index for the same product and return numbers that differ by four orders of magnitude. That is not a
Fingerprint or Not: Choosing the Right ZoomEye Query for a GeoServer Investigation
Two people can search an internet index for the same product and return numbers that differ by four orders of magnitude. That is not a data quality problem; it is a query design problem. The CISA advisory AA25-266A names one affected component, and turning that name into a defensible count requires deciding what relationship between asset and product the query is supposed to express.
Three relationships, three fields
An index like ZoomEye exposes several families of fields, and each answers a different question.
-
app=matches a product fingerprint derived from observable service behaviour. It answers "which assets behave like this application?" -
product=matches explicit product metadata reported by the service, which is far stricter. It answers "which assets state that they are this product?" -
title=matches text in the page title. It answers "which assets display this string to a visitor?" -
port=andservice=match network-level facts that are independent of the application. The results below show how much the choice matters for the same underlying product. | Query | Exact count | | --- | --- | |app="GeoServer"| 57,663 | |title="GeoServer"| 21,600 | |product="GeoServer"| 8 | |app="Apache Tomcat"| 581,630 | The spread from 57,663 to 8 is the central finding. A strict product field returns almost nothing, because most services do not volunteer structured product metadata. A behavioural fingerprint returns tens of thousands. Neither number is wrong; they describe different populations, and a report that quotes one without saying which is being used is not reproducible.
Prefer the fingerprint for discovery, and say so
For exposure measurement, the behavioural fingerprint is the right default. It captures assets that are functionally identifiable as the product regardless of what they declare. The trade-off is precision: a fingerprint match is an inference from observable behaviour and can include instances that have been heavily customised or fronted by a proxy.
That trade-off is manageable if it is stated. A sentence such as "57,663 assets matched the behavioural fingerprint for GeoServer on 2026-09-29" is accurate and useful. A sentence such as "57,663 GeoServer servers are vulnerable" is neither.
Combining fields to narrow intent
Filters are most useful when they encode a security hypothesis rather than a product name.
| Query | Exact count | What it isolates |
| --- | --- | --- |
| app="GeoServer" && service="http" | 47,664 | Endpoints reachable over plain HTTP |
| app="GeoServer" && port="8080" | 9,770 | Default listener, often untouched deployment defaults |
| app="GeoServer" && port="8443" | 625 | Alternate TLS listener |
| app="GeoServer" && port="8080" | 9,770 | default listener |
The service="http" result is the one with the clearest security reading. It isolates a large share of the population where traffic is not protected in transit, which matters because the reconnaissance and exploitation stages documented in the advisory both relied on ordinary HTTP requests.
CVE-scoped queries and their limits
A CVE field answers "which assets are indexed as affected by this CVE?" It is a discovery aid, and it is not equivalent to a version check.
| Query | Exact count |
| --- | --- |
| vul.cve="CVE-2024-36401" | 57,663 |
| app="GeoServer" | 57,663 |
The two totals match exactly. That coincidence is a property of how the CVE association is derived, and treating it as confirmation that all 57,663 instances run an affected release would be a mistake. The honest use of a CVE filter is to find candidate assets faster, then confirm version through your own inventory or an authorised check.
Conclusion
Query design is an editorial decision about what claim the number will support. For the GeoServer case, a behavioural fingerprint gives the broadest defensible exposure estimate, a title match gives a stricter subset, a product field gives almost nothing, and a port filter gives operational focus. Report which field produced the count, report when it was collected, and keep the claim narrower than the evidence.
References
- CISA, "CISA Shares Lessons Learned from an Incident Response Engagement," AA25-266A, September 23, 2025. https://www.cisa.gov/news-events/cybersecurity-advisories/aa25-266a
- ZoomEye search observations, exact-match counts, collected 2026-09-29 05:15 UTC. https://www.zoomeye.ai/
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.