I started collecting public web data for model pre-training at the end of last year. At first I went with some cheaper providers. Once concurrency went up, success rates dropped below 60%, and IPs were frequently flagged. After several rounds of switching, things gradually stabilized.
I tried pure residential, ISP, mobile, and mixed setups. Residential performed clearly better against Cloudflare, though traffic cost more. ISP worked better for tasks that needed session persistence. Mobile sometimes passed more easily in certain regions. I’m currently mainly using Helodata’s residential pool . City-level targeting is accurate and integration is straightforward, so no major failures lately.
Anyone else doing similar large-scale collection? What success rate are you able to maintain? Any particularly effective rotation or retry strategies? Would love to hear real numbers.
I tried pure residential, ISP, mobile, and mixed setups. Residential performed clearly better against Cloudflare, though traffic cost more. ISP worked better for tasks that needed session persistence. Mobile sometimes passed more easily in certain regions. I’m currently mainly using Helodata’s residential pool . City-level targeting is accurate and integration is straightforward, so no major failures lately.
Anyone else doing similar large-scale collection? What success rate are you able to maintain? Any particularly effective rotation or retry strategies? Would love to hear real numbers.










