Cet article est seulement disponible en anglais.
The IPIDEA network takedown by Google
On the 28th of January 2026, Google Threat intelligence published an article detailing the disruption of what it described as the world’s largest residential proxy network: IPIDEA1.
According to their findings, the entire ecosystem relied on the exploitation and deployment of malicious Software Development Kits (SDKs) within legitimate applications that enabled the transformation of devices into residential proxies servers, allowing third parties to route internet traffic through the devices’ residential Internet Protocol (IP) addresses. Access to these compromised devices was then resold as a commercial proxy service.

The IPIDEA proxy network was based on a network of multiple brands used to resell the proxies and Virtual Private Network (VPN) services:
- 360 Proxy (360proxy[.]com)
- 922 Proxy (922proxy[.]com)
- ABC Proxy (abcproxy[.]com)
- Cherry Proxy (cherryproxy[.]com)
- Door VPN (doorvpn[.]com)
- Galleon VPN (galleonvpn[.]com)
- IP 2 World (ip2world[.]com)
- Ipidea (ipidea[.]io)
- Luna Proxy (lunaproxy[.]com)
- PIA S5 Proxy (piaproxy[.]com)
- PY Proxy (pyproxy[.]com)
- Radish VPN (radishvpn[.]com)
- Tab Proxy (tabproxy[.]com)
According to their findings, these residential proxies providers relied on a few SDKs that were embedded in a large number of legitimate desktop and mobile applications, allowing the providers to turn users’ devices into nodes within their residential proxy networks:
- Castar SDK (castarsdk[.]com)
- Earn SDK (earnsdk[.]io)
- Hex SDK (hexsdk[.]com)
- Packet SDK (packetsdk[.]com)
The IPIDEA actors also rely on “free VPN” services which do indeed offer free services, but in exchange the device that connects to them also joins the network as a node. The following services were identified:
- Galleon VPN (galleonvpn[.]com)
- Radish VPN (radishvpn[.]com
- Aman VPN (defunct)
Finally, according to their analysis, they also identified infrastructures overlaps between several domains communicating with the SDKs suggesting that different SDKs may ultimately be operated by the same threat actor. Also, some of the domains communicating with the SDKs (what they call tiers-one Command-and-Control (C2) domains) were also observed as part of BadBox 2.0 botnet.
Why are we talking about all this?
Everything started with a scraping service. As part of our anti-scraping research and monitoring activities, we sometimes monitor web-scraping services and analyze their observable behavior to better understand how they operate. This includes studying behavioral patterns and client-side characteristics that may help fingerprint the service, as well as, where legally and technically permissible, identifying publicly observable elements of the infrastructure on which these services rely.

Coreclaw is a web-data extraction and web-scraping platform. In simple terms, it allows users to automatically collect information from public websites and turn it into structured data such as CSV, JSON, Microsoft Excel files, or API responses. The platform provides pre-built “workers” which are essentially ready-made for scraping and automation jobs. For example, it currently advertises workers for:
- Google Maps – extract business, contact details, review, opening hours etc...
- Google search – structured search result data (often called SERP2)
- Amazon – products, prices, reviews, rankings etc...
- Instagram / TikTok – public post/profile informations
- Youtube – channel and video informations

Using this service is quite simple, and it's possible to run the workers directly from the website or by using an API.
CoreClaw seems to be part of a HongKong based company3 named “Apex DataWorks Limited”:

However, the “/contact” URL references a different company “Vertex Digital CreationsLimited”, suggesting that this entity may be associated with the operation of the service:

The address remains consistent with Apex Data works: “Unit 9, 1/F, The Cloud, 111 Tung Chau Street, Tai Kok Tsui, Hong Kong”.
We found no reference to this exact company name. However, a company named “Vertex Apex Limited”4 was registered in Hong Kong 2 days before Apex Dataworks Limited was reportedly registered in Hong Kong on 24 of February 2026, two days before Apex Dataworks Limited
The address leads to almost nothing interesting as this is a corporate building where we can find many offices for a wide variety of activities (industry, technology, and even yoga).
We could go on at length with this type of research, but as it stands, it is not easy to establish a link between what appear to be registration addresses for what seems to be shell companies registered in Hong Kong and actual services.
Nevertheless, the combination of naming similarities, incorporation dates, and shared or closely associated registration details was sufficient to warrant further scrutiny.

What caught our attention
Our investigation could have ended there, but one option on CoreClaw caught our attention: the possibility to “select an execution node" (covering many countries worldwide). When tested, 129 locations were available as execution nodes.

The legitimate question that arises at this point is how does the service manage to provide such a large number of geographically distributed locations?
When using the service within a single country, we were surprised to see the number of different IP addresses they offer. Furthermore, sometimes the fingerprints changed slightly, potentially indicating the use of residential proxy networks.
Despite the main domain is “coreclaw[.]com”, we identified another URL that exposed the exact same website content:

At some point during the website loading process, an image is retrieved from an object-storage endpoint hosted at “oss.cafescraper.com”. We can also observed requests to “oss.coreclaw.com” which again confirms the link between the 2 domains:

The domain and main website are no longer accessible however, portions of the service’s documentation remain publicly available, mostly because one specific subdomain is still hosted through Mintlify and Vercel:

We then observed that the interface shown in the available screenshots is visually identical, or highly similar, to the CoreClaw interface. This provides an additional technical indication of a relationship between the two services, although visual similarity alone does not conclusively establish common ownership or operation...

Also, by default, the documentation is displayed in Chinese, as are some comments within the loaded JavaScript (JS) files, which potentially offers a clue as to the origin of the developers behind these solutions. However, they are not at this stage sufficient to reliably determine their nationality, geographic location, or origin.


Of course, both the domains were hosted on a common IP:
And we finally discover links to a GitHub account “CafeScraper”:


CafeScraper is part of an organization called “CoreClaw” on Github, the official account of the CoreClaw service.

Now, where it gets more interesting is that only two accounts list themselves as part of the CoreClaw organization on GitHub: “CafeScraper” already identified above, and another account that is quite well-known: Kael-Odin.
Kael-Odin is a technical writer and developer associated with web-scraping and proxy ecosystem. His public work focuses on technical subjects including residential proxies, anti-bot mechanisms, TLS fingerprinting, browser automation and web-scraping infrastructure. He also maintains scraping-related projects on Apify5, including a Crawl4AI-based web-content extractor and an Amazon search/product scraper.
Kael-Odin, in fact, is also the only listed developer of Thordata, a well-known residential proxy service:

Thordata currently advertises 100M+ residential IPs across 190+ countries. Its documentation describes the addresses as coming from “real residential networks” and supports country/city/ASN targeting, rotating sessions and sticky sessions.
In a February 2025 article6, he explicitly explains how Thordata fuel proxies:
- By paying customers directly providing access to their residential IPs
- By reselling residential IPs from partners
According to another blog post, he also claims to supply the network using an “opt-in SDK model”7. App developers embed a software development kit (SDK) into consumer apps, like free games or utilities. The final user agrees via terms of service to share their device's internet traffic and unused bandwidth in exchange for using the app.
Yes, we are indeed talking about the "terms of service", the long document that no one ever really reads. We aren't really at the same level of consent as the cookie opt-out windows mandated by the GDPR.

Finally, when a company develops both web-scraping solutions and the proxy infrastructure capable of supporting large-scale data collection, an obvious question arises: could those capabilities also be used to build and commercialize datasets directly?
We assume, for the purposes of this analysis, that the data was scraped legitimately and with the consent of the platforms involved...
At the time of writing, we identified 60 available datasets being offered. Among the largest ones were likely scraped from major social media platforms (TikTok, X, Youtube and Amazon) or from an AI provider (Anthropic).

In the sample we analyzed, the data from social media platform contains public datas scraped directly from user accounts. In the case of TikTok, for example, actual user images and videos URLs and profile data were scraped and are subsequently being resold.


The Claude opus data samples we downloaded contain likely intercepted (prompt, completion) pairs pulled from live claude-opus-4-8 traffic. This is the standard raw dataset shape that can be used for distilling a model on Claude's outputs or reverse-engineering its system prompts. Also, as it is possible to identify multiple different device_id and different topics (music scraper, care-rental scraping etc...), this dataset is likely aggregated traffic from people's private Claude Code sessions.
As for the question of how they gained access to this data, that is a very good question, but it is likely the access was not so legitimate (potentially the theft of users sessions cookies as it is very common?)

But the proxy network is ethically sourced. Right ...?
That's where things get interesting. Remember the previous infrastructure discovered? We have established a direct link between: coreclaw[.]com (the scraping service) and cafescraper[.]com (the previous/old project). We have shown that behind this scraping project there is also a developer who is part of the Thordata residential proxy network service.
Now where it gets interesting, taking a look at MX records for cafescraper[.]com, we identified this specific mail server hostname “mail.yougan.email”:

Pivoting on domains that answered with the “mail.yougan.email” MX record, yielded some quite interesting results:
- cherryproxy[.]net (instead of cherryproxy[.]com mentioned in the IPIDEA proxy network investigation by Google)
- castarsdk[.]com mentioned in the IPIDEA proxy network investigation by Google
- 360proxy[.]com mentioned in the IPIDEA proxy network investigation by Google
- datalabslmtd[.]com (a reference to “DATALABS LIMITED” cert that was mentioned in IPIDEA proxy network investigation by Google?)
So now we've reached the point where we're starting to wonder if there isn't a direct link between Thordata’s proxy network and the IPIDEA pool or other proxy providers.
Also, while looking at Thordata favicon, we identified 2 IPs where this favicon was identified in early January 2026:

The second IP is interesting as we can identify an overlap between 3 subdomains resolved on this IP:
- admin-api.thordata[.]net
- thor.worldrift[.]com
- thor-admin-api[.]worldrift.com

The domain “worldrift[.]com” was not mentioned in any report that we searched. We sought to understand the exact role of this domain, but without really managing to determine what it is used for. However, certificate transparency logs provide us with very clear indications regarding the infrastructure associated with this domain:
We also looked at passive DNS records, and the results (just a sample) speak for themselves. We can identify many residential proxy providers:
2 other domains were also seen using the Thordata favicon (3ab85cc6a6b88ce6c6)

Thordata x NetNut connection?
From one of the thordata[.]com subdomains, we also identified a direct reference to “NetNut”. This name might not ring a bell if you don't actively follow reports on residential proxy networks and other botnets.
NetNut is another commercial residential and ISP proxy provider. A recent publication from Synthient8 describes how NetNut also offers datasets and curated web-scraping services. This publication also provides evidence that NetNut’s proxy network is, at least in part, supplied through an SDK that turns consumer devices into proxy nodes (the “Popa” Android proxyware SDK).

Service observation results
Alright, so one question remains: is this proxy network truly legitimate or not?
That is open to debate : after all, if the user accepts the Terms of Service, the responsibility is arguably theirs to some extent. The real issue arises with proxy providers operating in a "gray area" that populate their networks using botnets and compromised devices.
We tried to answer that question by using the service. Our observations are based on a relatively small sample (9k+ distinct requests made from 2.6k IPs) but remain quite interesting.

We looked at the geographic distribution of the SpaceX / Starlink IPs we identified: Madagascar, Venezuela, Samoa, Kenya, Maldives etc... This is not the IP we were expecting when the service is heavily used by subscribers in the United States, Canada, United-Kingdom or Germany (based on external observations from electroiq9 as Starlink don’t publish official statistics about its consumer base).
As it is probably not economically viable at scale to buy Starlink services to resell proxy infrastructure, these IPs are likely not operator-owned infrastructure, but end-user devices recruited via proxyware SDKs, bundled apps (or malware ?).
Using Synthient lookup API10, we would then classify the observed IPs by type and identifying the proxy provider.
As we expected, most of the IPs are residentials which makes perfect sense for a service that also operates a residential proxy network:

Finally, we could also identify IPs classified as belonging to the IPIDEA pool (initial hypothesis) as well as those from other providers, notably BottingTool.

From a detection perspective, it’s worst mentioning some weird user-agents behaviors observed:
Conclusion
Unsurprisingly, our investigation identified strong correlations between scraping services, residential proxy networks, and data resale (specifically, data scraped using those proxy networks).
As the global race to train LLMs, which demands ever-increasing volumes of data, questions arise regarding the level of protection platforms offer against increasingly aggressive scraping actors. Detection can no longer rely solely on IP-based methods; instead, it must focus on behavioral analysis and client-side fingerprinting to distinguish genuine users from automated or potentially compromised clients. Furthermore, we remain highly skeptical of the "ethical" angle promoted by certain residential proxy providers. Indeed, infrastructure overlaps observed by many in the industry (Lumen, BitSight, Google...) suggest a close link between botnet networks and proxy providers, indicating that they all share, to varying degrees, the same IP pools.

The underlying risk exists at several levels:
- It has now been proven that there is a very strong correlation between residential proxy networks and a set of malware strains deployed specifically to achieve this goal and build a botnet. For instance, BitSight found a 15–26% overlap between the IP addresses of certain providers and malware such as Vo1d, Badbox, or RootSTV/Pandoraspear

- These residential proxy networks are not used solely for bypassing censorship in certain countries (Iran, Russia, or China) or for data scraping. Naturally, they are also constantly employed in various attacks and are particularly notable for their role in DDoS attacks. Notably, the Aisuru botnet was behind a series of record-smashing DDoS attacks in 202513.
- The central risk in any investigations resting on IP evidence is that a residential proxy inverts the meaning of the address: the IP you observe belongs to a victim, not a threat actor. In the 911 Socks5 case1415, the FBI describes how customers could commit cyberattacks, bomb threats, fraud, child exploitation, harassment, and export violations, knowing that the digital footprint would point back to the IP address of one of the botnet's victims. The scale makes clear this is not an edge case: over 19 million compromised IP addresses across more than 190 countries, including 613k in the United States alone. In our dataset for example, the problem compounds further: the Starlink addresses sit behind a carrier NAT, where a single public IP is shared among many subscribers while any one subscriber cycles through many IPs. So even a correct, timestamped ISP lookup cannot isolate a specific device.
The security recommendations regarding this issue are obvious:
- be wary of free apps and limit the number of applications installed on your devices (especially on Android, whether on mobile or smart TVs)
- If a product is free or very cheap (hi there AliExpress), it likely comes with a little gift designed to monetize traffic at the users' expense. Read the terms of service carefully (or use AI to help spot the clause explaining that your internet connection might be used as a relay point)
- Ensure your devices are up-to-date and do not expose your connected devices to the internet
Finally, I felt it was important to mention Pierluigi Vinciguerra’s blog on scraping.club16, which addresses the ethics surrounding proxy providers and the persistent issues associated with residential and mobile proxies. Regarding the FBI’s takedown of NetNut, he notably stated:
"How do you source your IPs, and can you prove consent? [...] What happens, both in contracts and technically, when abuse is found? If your provider can’t answer these questions, someone else will, maybe with a seizure banner."
References
Footnotes
1. https://cloud.google.com/blog/topics/threat-intelligence/disrupting-largest-residential-proxy-network
2. https://en.wikipedia.org/wiki/Search_engine_results_page
3. https://www.ltddir.com/companies/apex-dataworks-limited/
4. https://www.ltddir.com/companies/vertex-apex-limited/
5. https://apify.com/kael_odin
6. https://www.thordata.com/blog/residential-proxies/residential-ip-address
7. https://www.thordata.com/blog/residential-proxies/thordata-residential-proxy-the-ultimate-guide-to-ai-powered-data-collection-and-web-scraping-success
8. https://synthient.com/blog/popa-from-sourcing-to-distribution
9. https://electroiq.com/stats/starlink-statistics/
10. https://synthient.com/context/ip/{IP}
11. https://learn.microsoft.com/uk-ua/previous-versions/windows/desktop/legacy/dn904497(v=vs.85)
12. https://www.chromium.org/updates/ua-reduction/
13. https://krebsonsecurity.com/2025/10/aisuru-botnet-shifts-from-ddos-to-residential-proxies/
14. https://www.fbi.gov/news/podcasts/inside-the-fbi-podcast-the-911-s5-cyber-threat
15. https://www.ic3.gov/PSA/2024/PSA240529
16. https://www.scraping.club/p/when-the-fbi-knocks-on-your-proxy






