A Generalized Fingerprinting Framework for Deriving Searchable Features to Identify Publicly Exposed Infrastructure
Scanning of internet-exposed infrastructure has become a core methodology in
network security research, vulnerability assessment, and threat intelligence.
Threat intelligence analysts and researchers routinely scan the public internet to identify
exposed services, characterize device types, and infer software versions in
order to assess security posture and systemic risk. When a new vulnerability is disclosed,
national security teams face the task of identifying the hosts that run the vulnerable version.
Here, time and reliability are key, because vulnerable infrastructure has to be found before
attackers find it. Under NIS2, which significantly expands the number and
diversity of organizations under supervision, this task has become even more
important.
Therefore, this talk presents parts of the results of a master thesis performed at the request
of and in close collaboration with the National Cyber Security Centre (NCSC-NL).
It gives insights into a proposed methodology that guides threat intelligence
analysts to a fast and reliable search engine query, in this case Censys, in
order to identify vulnerable infrastructure. We present part of the methodology,
which is modeled as a decision tree whose leaves result either in an effective query
or in the conclusion that no suitable query can be derived. Alongside it, we present
the comprehensive feature table, in which features across domains were aggregated to
depict the most promising fingerprinting features. Certificates, exposed HTML code,
and even TTL values of OT devices can all be used as fingerprinting features to identify a specific
vulnerable product. Finally, we present parts of the prototype: a CLI-based
tool that allows researchers and experts to automatically identify suitable
features and the resulting query.
Description
### What this talk delivers
The talk starts with the problem of fingerprinting internet-exposed
infrastructure and why this task is necessary for CERTs and threat intelligence
analysts. Next, the methodology developed with NCSC-NL experts is presented. It
is modelled as a decision tree: by following its steps, the analyst arrives at a
real-world searchable query that returns the target product and version.
This query can then be used in search engines such as Shodan or Censys.
Alongside the methodology, the talk presents a table that is the result of
analyzing features such as certificates, exposed HTML code, URL paths, and TTL
values of OT devices across different literature and interviews, and aggregating
them into one final overview.
Finally, it is shown how parts of the methodology can be automated in the form of
a CLI-based tool that allows researchers and experts to automatically identify
suitable features and the resulting query. One or two practical results of the
tool are shown for a specific vulnerability, including the features the tool
generated for the query that identifies the affected infrastructure.
### The problem
The current state of the art is largely ad hoc and device-specific. In practice,
designing a scan for a new device or service requires manual reconnaissance and
target-specific scan logic. For example, when scanning a particular firewall or
network appliance, threat intelligence analysts must individually study its HTTP
responses, banners, TLS behavior, or protocol quirks. From these, they derive
fingerprints that uniquely identify the device and, in some cases, its firmware
version. Here, each analyst often follows their own path, risking missing
important features. This process is time-consuming, difficult to reproduce, and
poorly standardized.
Although research exists on which features are suitable for fingerprinting IoT
devices or web services, there is a scarcity of research focusing specifically on
a standardized methodology for how feature-based scans should be designed,
executed, or interpreted. Additionally, a lack of practicability can be observed
in terms of how to incorporate feature extraction based on a specific fingerprint
and apply it in the form of an effective query to existing scanning engines such
as Shodan or Censys. Consequently, different researchers and organizations often
perform scans in different ways, relying on implicit knowledge and personal
experience rather than shared, reusable practices.
### Why this matters for a national CERT
NCSC-NL is responsible for improving the digital resilience of the Netherlands,
with a particular focus on government organizations and NIS2-regulated entities.
To fulfill this role, NCSC-NL relies on infrastructure scanning to maintain
situational awareness of internet-exposed systems and to identify systemic risks.
Current scanning practices are largely device-specific and require manual
research for each new technology or appliance, and this approach does not scale
well under NIS2.
### Result 1 — observable features across domains
Multiple papers in the field of fingerprinting were analyzed to identify
observable features that are used for detecting publicly exposed infrastructure.
In addition, expert interviews were conducted in the field of threat intelligence
and analyzed to determine the features they use to identify specific products or
services in practice. The goal was to identify the most common features and merge
them into a summarized table that combines the findings from both the literature
and the interviews. Several feature families were identified, each grouping
similar features that share common characteristics: URL, path and routing;
headers and metadata; TLS certificates; resource artifacts; resource content;
network, transport and protocol; textual and semantic content; and technology
stack indicators.
Whereas existing research mostly focuses on specialized topics such as OT or web
service fingerprinting, this work provides an overview of common features that
apply across different domains. This is especially important because zero-days
are not domain-specific and affect OT systems, web services, routers, and VPN
systems — any system exposed to the internet. On OT devices, the product name and
version are often displayed in plain text on the login page. At the network
layer, Siemens PLCs and printers have different TOS, TTL and DF values; at the
transport layer the MSS of a Schneider PLC is 1380, while for Siemens it is 1460.
Based on the expert interviews, reliable results commonly rely on the following:
- HTML titles and HTML body content — high reliability when specific. Can
reveal product names or even versions in the HTML content of the login page. - URL paths and routing patterns — high reliability if available.
/admin.phpis generic and matches many services, while/dana-nais more
reliable because it is specific to Pulse Secure / Ivanti devices. - Certificate fields — reliable in specific cases. Vendor-signed
certificates may include product information in Issuer, Subject or SAN fields. - Combination of ports and banners — reliable only in combination.
- HTTP headers — low to medium reliability. May be altered by proxies,
gateways or load balancers. - Static resources, images, and favicons — supporting evidence.
### Result 2 — the methodology as a decision tree
The methodology was developed based on several qualitative interviews with senior
threat intelligence experts, as well as on published literature. These interviews
gave insights into current ad-hoc practices for extracting features to identify
exposed systems. By analyzing the interview results, formalizing the steps, and
finding common overlaps, an iterative process in the form of a decision tree was
constructed. The nodes represent decisions that must be made, and the leaves
represent the final outcomes: either a final query to identify exposed
infrastructure or the conclusion that no suitable query was found.
Start broad. The methodology starts with a high-level search, for example
"palo-alto", "SEL-3530", or "citrix netscaler". By starting with a broad
search term, the search engine can expose as many potentially relevant indexed
fields as possible. If an analyst starts directly with a specific query such as
vendor:"PaloAltoNetworks", useful features such as certificate fields or HTML
code might not appear and would be ignored.
Hunt for the product. After the high-level search, the next decision is
whether the scanning engine already provides structured product or vendor
metadata for the target. In the best case, a label already exists and aligns with
the hunt target, and it can be used to compose the query. For Sierra Wireless
devices running ALEOS, Censys exposes such metadata, for example
host.services.operating_systems.product: "aleos" and
host.services.software.vendor: "sierra_wireless". For the SEL-3530 Real-Time
Automation Controller there exists no specific label from Censys for product,
vendor, or operating system. The product-categorization path is therefore not
chosen; instead, the manual feature search path is followed, inspecting
certificate fields, HTML code, headers, or other indexed fields.
Hunt for the version. Hunting for the correct version is crucial because only
specific versions are usually affected by a vulnerability. The version field has
to match the corresponding product or operating-system field: searching for a
generic software version such as 5.0 would not be meaningful unless the scan
engine has first identified the product as pan-os, because otherwise the value
might refer to an unrelated product or service. This is why the software version
is always at least at depth 2 in the decision tree. In simple cases the version
is exposed directly in HTML, banners or headers; in more complex cases no version
term is mentioned anywhere, and artifacts that change between releases — CSS, JS,
favicon hashes, or login-page layout — have to be compared.
Verify before you trust. Even though Shodan or Censys assign a categorization
to products or versions, a service or device can still be categorized incorrectly
if the underlying features are too broad. Search engine results should therefore
not be trusted blindly, but always verified against reliable information such as
official documentation, observed login pages, certificates, or service banners.
There exists no ground truth: it is usually not possible to contact the providers
of the returned infrastructure and verify whether they are actually running a
SEL-3530 on a specific IP address. The proposed confidence level is derived from
an evaluation of the verification steps, as a way to add transparency rather than
as a formal research result.
### Two worked examples
Sierra Wireless ALEOS — the labelled case. Censys has already fingerprinted
the devices internally and exposes structured metadata in several fields. The
challenge is that the analyst has to know the relationship between acemanager
and ALEOS: ACEmanager is the web-based management interface used by Sierra
Wireless AirLink devices running ALEOS. The metadata is also spread across
different Censys fields, which is another reason why the methodology starts with
a broad search.
SEL-3530 RTAC — the unlabelled case. No usable product or version label
exists, so the fingerprint is built from matched fields such as
web.cert.names: "SEL-3530 RTAC" and web.hardware.product: "sel-3530". A
returned host can fall into one of two cases: the host actually runs a SEL-3530
RTAC, or an unrelated website merely contains the text SEL-3530 in its HTML
title, which makes the result a false positive. The resulting query is therefore
based only on the certificate feature while excluding results that only match in
the header.
### Result 3 — automating parts of the methodology
A prototype was created to investigate if and how automation can support feature
extraction and scan design. It is a CLI-based tool that allows researchers and
experts to automatically create a query, and the whole prototype is based on and
derived from the designed methodology. Two practical runs are shown: ALEOS, where
Censys already exposes product metadata, and SEL-3530, where it does not.
The methodology was first simplified for automation. Product and version
identification were combined into a single step, because following every layer of
the original tree triggered too many Censys calls. Every run then branches into
one of two pipelines. Pipeline A checks whether aligned product and version
metadata is already available and, if so, composes the query from those fields.
Pipeline B is triggered when no aligned metadata exists: suitable features
such as certificates, HTML, or headers then have to be identified from the
returned host records.
To allow for better flexibility and enhanced contextual understanding, the
prototype was extended with an optional agent architecture (--agent). As LLMs
recently showed great performance in contextual understanding of text, agents
were added because unrelated, or even completely different, keywords sometimes
appear in the results that are relevant for the final query but only have a
contextual relation to the searched product. A parser-only approach cannot
recognise that Palo Alto, PAN-OS and GlobalProtect belong together, while
linux, nginx or openssh do not. The context given to the agents via
--context is usually the advisory text of a vulnerability. OpenAI GPT-4o mini
was used as the underlying model; Pydantic is used to enforce a fixed object
structure per agent, because LLMs often make mistakes in the JSON they return
and this would make the whole pipeline fail.
Giving all bucket results to a single agent would be straightforward, but it
would violate separation of duties. The agent architecture therefore splits the
problem. For Pipeline A:
- Deterministic bootstrap — four Censys facet aggregate calls on the broad
query. - Agent 1 (bucket filter) — classifies every bucket value as
keepor
dropwith a reason, using the advisory as context. Generic stack values
(linux,nginx,openssh) and versions outside the vulnerable range are
dropped. - Agent 2 (query composition) — one LLM call; returns a CPQL query from the
kept buckets only, or"FALSE"if no suitable features are present, which
triggers Pipeline B.
If Pipeline A cannot find a usable label, Pipeline B takes over. A predefined
parser is replaced by agents that work directly on the returned host sample:
- Deterministic bootstrap — one broad full-text search on the quoted
identifier (page size 50). Unlike Pipeline A there are no facet aggregate
calls; the agents work on the host records themselves. - Agent 3 (manual feature extraction) — one LLM call per host; extracts
unique fingerprint features (certificates, HTML, headers) and keeps them if
they uniquely identify the product. - Agent 4 (aggregation) — aggregates the per-host features into strong
shared indicators that appear across many hosts. - Agent 5 (query composition) — one LLM call; returns a CPQL query from
those indicators.
This separation improves transparency and root-cause analysis, because each
agent has a single responsibility and can be tested independently.
The limits are presented as clearly as the results. A clear bottleneck was the
limited number of Censys API calls that could be made, which limited agent-based
trial-and-error. The automation mainly focused on Censys, and some parts, such as
querying buckets, might not work with other search engines such as Shodan. The
agent variant is also non-deterministic: in some runs not all previously
identified versions were chosen.
What you can take home
This work supports researchers and threat intelligence teams with a practical
guide for identifying usable fingerprinting features in a more transparent,
reproducible, and reliable way. It is also practical, as it incorporates common
search engine tools such as Shodan and Censys. The work shows many possible ways
to fingerprint a device, and lists usable features across domains — certificates,
CSS characteristics, HTML titles, URL paths, and protocol traits — that can be
translated into a query to identify a specific product and version.
Furthermore the same features are not only useful for finding vulnerable infrastructure.
They can equally be used to fingerprint attacker infrastructure, or any other exposed service an analyst needs to identify. That is also why minor information leakage, through login pages,
certificates, URLs, or version strings, can reveal sensitive information about a
company's systems. For OT products, the product and version are often already
exposed on the login page. Vendors should ask whether their login page really has
to expose the product or version, whether the certificate really has to contain
product information, or whether CSS and HTML files can be made more generic.
The research also showed how the process of identifying features that can be used
to find exposed vulnerable devices can be automated. Agents can be pipelined to
extend the tool with contextual understanding.
Coming from a software development background and having transitioned into the security domain, I really enjoy building prototypes, experimenting with new ideas, and exploring areas such as threat intelligence.
Bob is a senior (as in old and working for an eternity) CTI-Specialist of the NCSC-NL.