Skip to content
Cybersecurity

Cyber Warfare Research: Using Internet Scan Data (2026)

/ 12 min read / Malik Tanveer Dhool

How threat-intelligence analysts use internet-wide scan data to study cyber warfare: certificate pivoting, infrastructure baselining, reuse tracking across threat reports, and the hard limits of what scan data can prove. Real case studies from Censys research and documented investigations.

← Back to Blog
Share X in

WarBrief Live | October 6, 2026 | Cyber Intelligence

When modern states fight through the internet, their tools leave traces on it. Command servers, spyware infrastructure, and vulnerable devices all sit on public IP addresses that can be found, counted, and studied from anywhere. That is why Censys cyber warfare research has become one of the most practical disciplines in threat intelligence: analysts use internet-wide scan data — records of what every reachable device on the internet looks like, right now and months in the past — to study how adversaries build and reuse their infrastructure. This article explains the craft itself: the methods researchers use, real cases where the methods paid off, and the hard limits of what scan data can ever prove.

Key Takeaways

  • Analysts study cyber warfare with internet scan data through four repeatable methods: baselining exposed infrastructure, tracking reuse across threat reports, pivoting on certificates, and correlating scan data with intelligence reporting.
  • Citizen Lab mapped the spyware vendor Candiru’s command infrastructure by pivoting between Censys host and certificate records — a case that led to the discovery of two undisclosed vulnerabilities.
  • Published investigations of the PolarEdge botnet, the Notepad++ hijack, and Remcos command servers show how historical scan snapshots reveal infrastructure other tools cannot reach back in time to see.
  • Scan data shows what is visible on the internet — not who runs it. Similarity between servers is a lead, not an attribution, and exposed does not mean compromised.

What Is Internet-Wide Scan Data, and Why Do Cyber Warfare Researchers Use It?

Internet-wide scanning means probing every reachable address on the internet and recording what answers: which ports are open, which services run on them, which certificates they present, and how all of that changes over time. Censys’s platform scans across all 65,535 ports and more than 200 protocols, and it keeps historical snapshots of every internet-connected asset — so a researcher can look at a suspicious server today and see what it looked like six months ago.

For people studying nation-state cyber operations, this matters. Malware analysts get samples; network defenders get logs. But only scan data gives an independent, global picture of the infrastructure itself: the servers campaigns rely on before they are taken down, the certificates operators reuse out of habit, and the vulnerable devices sitting in a conflict’s blast radius. Censys formalized its research function in March 2026 with the launch of the Censys Advanced Research Collective (ARC), whose stated mission includes linking malicious infrastructure to nation-state actors and tracking exposed devices in relation to international conflicts. If you want the basics of how Censys collects this data, our explainer on Censys Inspect covers the scanner itself CensysInspect Explained: Internet Intelligence Guide.

How Do Analysts Actually Use Internet Scan Data in Censys Cyber Warfare Research?

AI-generated illustration

Four methods recur in real investigations. None require access to the adversary’s machines — only to the public internet’s memory.

1. Baselining exposed infrastructure

The simplest method is also the most powerful: count everything that is visible. In mid-2026, Dutch intelligence agencies reported a Russian state espionage campaign that was quietly hijacking insecure IP cameras across Europe and Ukraine to monitor weapons deliveries and troop movements. Researchers used Censys scan data to quantify the target surface: more than 87,000 exploitable internet-connected cameras across EU, NATO and Ukrainian territory, including over 4,000 in Ukraine and nearly 2,000 in the Netherlands. Many ran software with years-old flaws — including a Dropbear SSH issue known since 2017 — and 541 camera services were exposed through vulnerabilities known to be exploited in the wild.

Baselining turns a warning into a measurable risk. The same approach appeared in April 2026, when the FBI, CISA and the NSA warned that Iran-linked threat actors were exploiting internet-exposed Rockwell Automation PLCs in attacks on operational technology. Censys researchers counted 5,219 such devices online — 74.6% of them in the United States, many on cellular networks — and their analysis suggested multiple exposed IPs were tied back to a single compromised engineering workstation, expanding the attack surface beyond the agencies’ initial disclosure. Baselining cannot prove which devices were attacked, but it tells defenders exactly where to look first. For more on how scan data relates to national infrastructure exposure, see Censys and Critical Infrastructure: Finding Exposed Systems and Why Internet Intelligence Matters for Military Cybersecurity.

2. Certificate pivoting

If baselining is counting, pivoting is hunting. Every server that uses TLS encryption presents a certificate, and operators — even careful ones — leave fingerprints in them: a reused subject name, a distinctive self-signed issuer, an odd certificate chain. Find one suspicious certificate, search for every other host presenting it, and a single server becomes a whole cluster.

The landmark public example is Citizen Lab’s investigation of Candiru, a private-sector offensive actor selling spyware to governments. As recounted on the Censys research blog, researcher Bill Marczak’s team used Censys’s IPv4 host and certificate datasets to map Candiru’s command-and-control infrastructure. They started with a self-signed certificate tied to “candirusecurity[.]com” — known from a 2015 corporate registration filing — then pivoted iteratively between hosts and certificates with historical lookback, surfacing certificates for more than 750 websites that Candiru spyware infrastructure was impersonating. As Marczak put it:

“We were curious about mapping out command and control infrastructure — IPs, domains, certificates — with the ultimate goal of understanding Candiru’s global footprint.” — Bill Marczak, Senior Research Fellow, Citizen Lab (via the Censys research blog)

The investigation did not end at mapping. Citizen Lab recovered the spyware sample and passed it to Microsoft — which found two previously undisclosed vulnerabilities and identified more than 100 other targets among human rights defenders, journalists, activists, and politicians. Certificate pivoting turned one data point into an international disclosure.

3. Tracking infrastructure reuse across published threat reports

The third method starts where a threat-intel report ends. When vendors like Sekoia, Rapid7, or Microsoft publish an investigation, they name infrastructure: an IP, a domain, a malware family. Scan-data analysts pick up those indicators and ask what else the same operators built — including infrastructure the original report never saw.

A clean example comes from Censys’s 2025 State of the Internet Report. In February 2025, Sekoia researchers uncovered PolarEdge, an IoT botnet active since at least late 2023 that had exploited CVE-2023-20118 in Cisco Small Business routers to plant a custom Mbed TLS backdoor. Censys’s report spotlights how its researchers took Sekoia’s identified payload-distribution host — 119.8.186[.]227, on Huawei Cloud in Singapore — and pivoted through scan data, reaching back to historical snapshots from February 11, 2025, around the time of the attacker’s activity, to examine the host’s services and certificates and find adjacent infrastructure.

The same pattern played out in Censys’s February 2026 analysis of the Notepad++ supply-chain hijack, in which attackers redirected the text editor’s update traffic to servers distributing the previously unknown Chrysalis backdoor — an operation Rapid7 attributed to the China-linked APT Lotus Blossom (also known as Billbug). Censys found the Chrysalis command-and-control certificate for api.skycloudcenter[.]com reused across two hosts, and a Cobalt-Strike-like certificate reused across multiple hosts — reuse patterns that enabled infrastructure pivots. It also reconstructed the timeline of a staging host that cycled through SSH-only, short-lived HTTP/HTTPS on non-standard ports, and brief VPN exposure from February 2025 through January 2026, reading the churn as a reusable staging asset rather than a single-purpose server.

4. Correlating scan data with threat-intel reporting

The fourth method is synthesis: scan data rarely answers “who” by itself, but stacked against malware analysis, incident reporting, and victim telemetry it becomes corroborating evidence. In the Notepad++ case, Rapid7’s endpoint analysis identified the backdoor and the likely actor; Censys’s scan-data analysis identified the infrastructure behavior and applied the brakes where the evidence fell short — noting, for example, that a changed SSH host key on the staging host could indicate a change of ownership rather than continuity.

Censys’s ArcaneDoor infrastructure analysis, also published in 2026, shows the same discipline from the other side. Researchers pivoted on certificate issuer names found on suspect hosts — including an unusual “Gozargah” certificate — and traced the term to an open-source censorship-circumvention project. They framed the result exactly as it should be framed: the analysis “suggests potential ties” to a Chinese-based actor, not proof of them. Correlating scan data with reporting is where investigations get honest: every pivot adds weight, and every mismatch gets named.

Case Study: Counting a Botnet’s Footprint at Internet Scale

Not every investigation ends in attribution. Some of the most useful cyber warfare research is descriptive: how big is this thing, where does it live, how is it changing?

Between October 14 and November 14, 2025, Censys tracked more than 150 active Remcos command-and-control servers — Remcos being a commodity remote access trojan used by many operators. Most servers listened on port 2404, Remcos’s default, with additional activity on ports 5000, 5060, 5061, 8268, and 8808. Hosting concentrated in the United States, the Netherlands, and Germany, with a significant share on inexpensive, lightly vetted providers such as COLOCROSSING, RAILNET, and CONTABO. Certificates were sometimes reused across multiple IPs — a sign of template-based configuration and minimal obfuscation — which made cluster linkage straightforward.

Why does this matter for studying cyber warfare? Commodity tools like Remcos are the background noise of state-sponsored operations: cheap, disposable, deniable. Tracking a malware family’s footprint over time — its persistence, turnover, hosting preferences — builds the baseline against which the next sophisticated campaign stands out.

What Can’t Internet Scan Data Tell Us?

AI-generated illustration

Scan data is extraordinarily powerful — and researchers who use it well are explicit about its limits.

1. Exposed does not mean compromised. A scan can prove a camera is reachable and vulnerable, not that anyone broke into it. In the Russian IP-camera reporting, the exposure numbers came from scan data; the compromises came from intelligence agencies. Confusing the two exaggerates every finding.

2. Similarity is a lead, not an attribution. Two servers presenting the same certificate or TLS fingerprint are probably related — but shared tooling, templates, and resale of cloud infrastructure can produce identical artifacts from unrelated operators. Censys’s own Cobalt Strike methodology warns that suspicious certificate characteristics are not proof without confirmed configuration.

3. Patch status is often invisible. Scans see service banners and handshake details, but as Censys researcher Himaja Motheram told TechTarget in 2023, “patch details aren’t always visible to Censys’ passive scanners.” A vulnerable-looking service may be patched behind the banner.

4. Scans see the outside, never the inside. Scan data shows services a host exposes. It cannot see what malware runs on it, who logs in, or what data it holds. That is why correlation with endpoint telemetry and threat reports is a method in itself, not an afterthought.

5. Hosts change hands. IP addresses are reassigned, cloud instances are resold, and SSH host keys change. A server’s history is a timeline of observations — attributing every observation to one operator requires evidence beyond the timeline itself.

The practical rule analysts follow: scan data is excellent at generating hypotheses and terrible at confirming them alone. Every conclusion it produces should carry its confidence level, and every “suggests” should be one word, not three.

Frequently asked questions

How do threat-intelligence analysts use Censys scan data to study cyber warfare?
They use four repeatable methods: baselining the population of exposed infrastructure in a region or sector, pivoting on reused TLS certificates to find related servers, tracking how infrastructure named in published threat reports evolves over time, and correlating scan observations with malware analysis and intelligence reporting to build or check hypotheses about campaigns.

What is certificate pivoting in threat-intelligence research?
Certificate pivoting means starting from one suspicious TLS certificate — a reused subject name, a self-signed issuer, an unusual chain — and searching internet-wide scan data for every other host presenting the same certificate. Citizen Lab used this method to map the spyware vendor Candiru’s command infrastructure, surfacing certificates for more than 750 impersonated websites from a single self-signed certificate.

Can internet scan data prove who is behind a cyber attack?
No. Scan data shows what is visible on the internet — open ports, services, certificates, and their history. It can establish that infrastructure is related (shared certificates, templates, timelines), but attribution to a specific actor requires additional evidence such as malware analysis, victim telemetry, or intelligence reporting. Researchers frame scan-data findings as leads and correlations, not proof.

What is the difference between exposed infrastructure and compromised infrastructure?
Exposed infrastructure is reachable from the internet and may be vulnerable; scan data can measure it directly. Compromised infrastructure has actually been broken into and used by an attacker, which scan data cannot confirm on its own. In the 2026 Russian IP-camera espionage reporting, for example, scan data quantified tens of thousands of exposed cameras while the compromises were reported by intelligence agencies.

Where can I learn more about the tools behind this research?
See CensysInspect Explained: Internet Intelligence Guide (what Censys Inspect is and how scanning works), Censys vs Shodan for Threat Hunting Teams Compared (2026) (Censys vs Shodan for threat hunting), and How Internet Intelligence Helps Investigate Cyber Attacks (a practical guide to investigating cyber attacks with scan data).

Conclusion

Cyber warfare research with internet scan data is, at its core, an observational science. The methods — baselining, certificate pivoting, reuse tracking, and correlation — do not hack back, exploit, or participate in attacks. They watch, count, compare, and remember: one certificate led to a spyware vendor’s global footprint, Sekoia’s PolarEdge indicators were extended through historical snapshots, and thousands of exposed cameras and controllers were counted against real espionage warnings.

But evidence has grades, and honest researchers label them. Scan data can prove what the internet exposes; it cannot prove who is behind it, whether a vulnerable device was ever touched, or what happens inside a server observed only from the outside. The analysts who use Censys data best say “suggests” when they mean “suggests” — and their conclusions survive because they did.

But evidence has grades, and honest researchers label them. Scan data can prove what the internet exposes; it cannot prove who is behind it, whether a vulnerable device was ever touched, or what happens inside a server it can only observe from the outside. The analysts who use Censys data best are the ones who say “suggests” when they mean “suggests” — and whose conclusions survive because they did.

Sources and Further Reading

Written by

Malik Tanveer Dhool

Defense and intelligence analysis for WarBrief.live. Covering conflict, technology, and geopolitical strategy.