Destination Attribution
discover-process-egress reports which external destinations a process talks to. Those destinations are usually bare IPs, so the tool attributes them to a provider — this page explains how far that attribution can be trusted, and why it stops where it does.
Why not FQDNs
The obvious answer would be hostnames. The PCE only populates dst.fqdn when it has DNS visibility for a flow, and for endpoint egress it usually does not. That is a data-availability limit in the flow record, not something the server can work around.
Two approaches were measured and rejected:
Reverse DNS is actively wrong. 160.79.104.100 — a real Anthropic address that Claude Desktop connects to — has the PTR record fatpipe.juilliard.edu. A report built on rDNS would state that Claude was talking to Juilliard. Confidently wrong is worse than a bare IP.
Live whois adds nothing for the ambiguous cases. ChatGPT egress lands on Cloudflare (172.66.0.243, 162.159.140.245), and whois returns Cloudflare, Inc. — exactly what the shipped table already says. The ambiguity is structural, not a data gap: TLS SNI is what distinguishes one Cloudflare tenant from another, and Illumio never sees it.
IP-to-organisation is the honest ceiling from flow data alone. For true hostname resolution, correlate with DNS server or proxy logs.
What it can identify
21 providers, 3,412 ranges, all from each vendor’s own published list or from RDAP. Lookup is longest-prefix, so a specific service range beats the broad cloud range containing it.
| Provider | Confidence | Ranges | Source |
|---|---|---|---|
anthropic | likely | 1 | rdap |
atlassian | likely | 166 | published |
box | likely | 1 | rdap |
dropbox | likely | 1 | rdap |
github | likely | 108 | published |
salesforce | likely | 61 | published |
workday | likely | 1 | rdap |
zoom | likely | 49 | published |
aws-api_gateway | ambiguous | 214 | published |
aws-cloudfront | ambiguous | 243 | published |
aws-route53_healthchecks | ambiguous | 57 | published |
aws-s3 | ambiguous | 1129 | published |
azure-hosted | ambiguous | 4 | curated:coarse-cloud |
cloudflare-fronted | ambiguous | 22 | published |
fastly | ambiguous | 21 | published |
google-cloud | ambiguous | 1103 | published |
google-services | ambiguous | 138 | published |
microsoft365 | ambiguous | 40 | published |
microsoft365-exchange | ambiguous | 34 | published |
microsoft365-sharepoint | ambiguous | 10 | published |
microsoft365-skype | ambiguous | 9 | published |
Verified by resolving each vendor’s well-known hostname and classifying the result:
| Destination | Reports as |
|---|---|
gmail.com, drive.google.com | google-services |
outlook.office365.com | microsoft365-exchange |
login.salesforce.com | salesforce |
zoom.us | zoom |
github.com | github |
www.myworkday.com | workday |
What it cannot identify, and why
Some SaaS cannot be attributed by IP at all, because they do not own the addresses they answer on. Checked by RDAP:
| SaaS | Resolves into | So it reports as |
|---|---|---|
| Slack | Amazon | aws-* or nothing |
| Zendesk | Cloudflare | cloudflare-fronted |
| DocuSign | Microsoft | azure-hosted |
| ServiceNow | Akamai | nothing |
This is a property of their hosting, not a gap in the table. No range file can fix it: the address belongs to the CDN, and thousands of unrelated tenants answer on the same one. Distinguishing them needs TLS SNI or DNS logs, which Illumio flow data does not carry.
So an empty result means “not attributable”, never “no traffic”.
Confidence tiers
Where the data comes from
The server performs no network I/O for attribution. It reads src/illumio_mcp/data/ip_ranges.json and does offline CIDR containment.
That matters for three reasons: the server stays deterministic and testable; it deploys into restricted-egress environments; and customer destination IPs are never sent to a third-party enrichment service — that is the customer’s egress topology, which is the sensitive thing the tool exists to analyse.
The table is regenerated by scripts/refresh_ip_ranges.py, which runs in CI only (.github/workflows/refresh-ip-ranges.yml, weekly). Sources:
- RDAP for vendor-owned space — the registry’s own netblock and registrant. If a range changes hands, the refresh fails loudly rather than continuing to attribute an unrelated company’s traffic to an AI vendor.
- Vendor-published lists for shared infrastructure that publishes them (Cloudflare).
- Curated coarse ranges for cloud providers, kept deliberately imprecise because the claim they support is only “hosted on this cloud”.
Changes arrive as a pull request, never a direct push: a range change alters what the tool reports about customer traffic, so it is reviewed like any other behaviour change.
Running it manually
python scripts/refresh_ip_ranges.py # regenerate
python scripts/refresh_ip_ranges.py --check # verify currency, write nothing
Exit 1 means it wrote a valid table but hit something worth a look — a fetch failure, or a range that changed registrant.
Known limitation
A static table rots. Providers re-allocate ranges with no announcement, and nothing at runtime detects it. The weekly refresh is what bounds that drift; the generated_at field in the data file records when the shipped table was last confirmed.
Tests assert that every CIDR parses (a typo silently disables a provider) and that ranges do not overlap across providers (overlap makes attribution depend on table order). Neither can verify the ranges are still current — only the refresh does that.
The first refresh run caught two errors in the hand-written table it replaced: the Anthropic entry was /23 where ARIN says /21 — missing real Claude traffic — and the openai mislabel described above.