Rendered at 08:13:57 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
baud9600 2 hours ago [-]
One thing I’m wondering is whether we need an open source set of technical patterns and libraries for dealing with this changing traffic mix.
Millions of small sites and creators don’t have the ability to design their own protections against large scale automated access. If useful content now attracts aggressive crawler traffic, many sites will be too expensive or unreliable to run.
Is part of the answer a community response? Perhaps a community-maintained toolkit, based on traffic data, that host sites can apply? It could include standard agent identification, rate limiting, traffic classification, access policies, caching, challenge mechanisms, logging, attribution and usage control etc etc.
In effect, we need much stronger road rules for today’s automated traffic, available as open technical patterns and libraries rather than every site owner having to invent this alone (they won’t).
m-i-l 11 minutes ago [-]
Yes, a community effort against the botnets would be great.
The "good bots" which identify themselves and follow the rules are easy enough to block, so not a problem. It is the "bad bots" which pretend to be real users and hide behind residential proxies, and so are almost impossible to block at the moment, which are the problem.
Given that there are big companies openly (i.e. on the clearweb, not even darkweb) selling access to these residential proxy botnets of compromised smart TVs[0] and mobile phones and other devices, can we not simply get a database of the these IPs and block access from them? I would venture that almost every single one of those residential users are unaware that they have devices in their homes which have been compromised and are being abused in this way, so if they were to start seeing messages from more and more sites along the lines of "Access to this site has been blocked because unusual traffic has been detected from your computer network. Please check all devices on your network and remove any malware which may be routing this traffic." then maybe we could start addressing the problem at the source.
The community response was respecting robots.txt, but since that was more of a gentleman's agreement than a legal requirement, hungry AI scrapers disregarded it.
Which brings us to the old fashioned flood control mechanisms. That is, the toolkit you propose already exists and has been used for decades in various iterations to protect against various forms of attack (slashdotting, DDOS attacks, overzealous search engines, and now AI scrapers).
Have you looked into those before? Companies like Cloudflare have been at the forefront of this field for a long time now.
abetusk 14 hours ago [-]
I think people are missing one of the points of the article. It's not just that agents are hammering the site, it's that there might be lurking vulnerabilities that allow malicious usage, which is why it went down, then came back up with a fraction of the data and a reduced design.
The article says (speculates?) that malicious users are trying to get privileged access for an edge in prediction market betting. From the article:
> If you could see The Numbers data before everyone else, every single week, you would have a significant edge over all the other traders - learning the answers slightly ahead of publication would allow you to front-run the trades.
squidproquo 6 hours ago [-]
If it's the case that prediction market traders were trying to front-run the market outcomes, the prediction markets should put the fix on their end by locking down trading before a market closes with some time buffer to prevent this.
eru 2 hours ago [-]
Why should that be necessary?
You can unilaterally stop trading 'before a market closes with some time buffer to prevent this.' No need for centralised action.
Cthulhu_ 15 minutes ago [-]
That's assuming the prediction markets all play fair and by the same rules. If Polymarket were to do that, its competition could choose not to. Since these exchanges are not tied to any one country (or jurisdiction), its users would jump ship. Because the users don't care.
See also stock market trading, where companies would go to great lengths to shave nanoseconds off of getting information and putting in trades. Now apply this min/maxing to worldwide, unregulated and anonymous.
This is what tech libertarians / cryptobros want.
blurgo 5 hours ago [-]
[dead]
podgietaru 14 hours ago [-]
"One Reddit theory even suggested it was a deliberate rug pull designed to cripple the free site to push people towards paid products."
I know this isn't really the point of the article, but I've been thinking about this a lot. I wonder if we're going to see more resources go this way.
I used to publish little doo-dads as opensource software. Not because it was something that was legitimately ground breaking or anything (it absolutely wasn't) but because I had a problem, and I thought "heh wouldn't it be cool if someone else had a similar problem and could use my resource for it."
But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.
And I can see why maintainers of sites like this, or other free but incredibly useful resources might start to get irate at that.
chrka 2 hours ago [-]
I also recently made an open-source project with 200 GitHub stars private. I never had a problem with others using it as the basis for their own projects. In fact, that happened, and I received credit for it. But LLMs just hoover up everything, process it, and then spit it back out as if it were their own.
II2II 13 hours ago [-]
> But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.
The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem.
> And I can see why maintainers of sites like this, or other free but incredibly useful resources might start to get irate at that.
Those who object to the scraping fall into several camps, but the biggest complaint I am hearing is that it increases both maintenance costs and time. In other words: it sucks when people are using your work in a manner that you find offensive, but it goes beyond that by doing actual harm.
TeMPOraL 25 minutes ago [-]
> The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem.
> In other words: it sucks when people are using your work in a manner that you find offensive
That to me still reeks of Dog in the Manger mentality. If you publish something for the world to use, you should neither care nor even track, much less discriminate by (or suddenly seek compensation for) who is using it.
Cthulhu_ 12 minutes ago [-]
Agreed, this is what you agreed to (implicitly or explicitly) when uploading stuff to the internet. In fact, open source embraces this (hence the 'open').
But people are free to not publish things or post things online with a more restrictive license. Not that a license stops things from being indexed.
joshmarinacci 13 hours ago [-]
A change in quantity can become a change in quality.
rjtavares 2 hours ago [-]
I think the toxicology proverb applies here perfectly: The dose makes the poison.
dotancohen 1 hours ago [-]
I believe that it was Stalin who phrased it most eloquently: Quantity has a quality all its own.
antisthenes 12 hours ago [-]
> The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem.
People doing it with a couple of machines and residential proxies versus Anthropic doing it with 2 data centers worth of machines.
Scale matters.
weitendorf 13 hours ago [-]
I open source as much software as I can because I want the models to train on it and get better at it!
Scoundreller 13 hours ago [-]
I feel the same about my old blog articles.
I used to rank pretty well trashing crappy credit cards and encouraging people to switch to better options, then Google decided that 10 results for the card issuers website was better.
Same shit when I manually wrote proto-gethuman posts on calling telecoms/banks/etc (and also pushed visitors to try an Indy ISP or credit unions), then Google felt it was better to drive users to the telecom’s website that wants you to do anything but call them.
Please do train on my pre-LLM gold!
da02 12 hours ago [-]
Do you have any active sites or social media with your recommendations?
(I never publish my recommendations because they seem to complicated for people. Like using 30 GB plan/$10/monthly from T-Mobile for data and then using Tello for $10 plan for voice/text. This would require a 2 esim/sim card phone. I am currently using a Moto G Power 2024 phone from eBay $90/new, which was better than the $200 slightly used Pixel 6a from Swappa. Most people would just save the hassle, get a Galaxy phone with a phone contract.)
Scoundreller 10 hours ago [-]
Nothing active anymore, all converted to static sites.
If you’re on page 2 of the results, you effectively don’t exist so I stopped bothering/benefiting from display ads.
Coincidentally, I did try to get deeper into the cellular service side (it’s another high margin and high customer value segment), but I did better on the finance side.
ryukoposting 7 hours ago [-]
> Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.
I suspect the next AGPL (if not normal GPL) will explicitly prohibit training of closed models against the source material. Ideally it will prohibit training of open weight models too. Either share the entire process, or go piss up a rope.
dotancohen 1 hours ago [-]
My understanding of the GPL is that it currently prohibits the training of LLMs, no need to explicitly mention them. MIT does allow it, however.
chrka 7 minutes ago [-]
Exactly. Just like using LibGen is prohibited.
protocolture 2 hours ago [-]
>But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.
I dont get this.
You provided something for free to help people, but dont want to do that anymore because it might go into training data and help many many more people?
So far LLMs have been loan funded donations of loss leading services. They might never actually make their first dollar of profit.
Meanwhile ISPs the world over have been monetizing access to your content.
I feel like what you mean is that you want control over attribution.
vachina 13 hours ago [-]
> But now I'm really reluctant to give more stuff to the free web.
I feel the same way too. But guess what, all the code you did not publish gets into the training corpus anyway (when you gave Codex or whatever full read access to your filesystem).
bigstrat2003 12 hours ago [-]
That's only if one is stupid enough to give an LLM access to their computer. Don't do that, it's incredibly irresponsible security wise.
dotancohen 1 hours ago [-]
The vast majority of people are in fact this stupid.
I recently had trouble connecting to a new AP on my laptop. After a quarter hour of frustration, I connected to the AP of my phone, asked Claude Code what the problem is, and she found the issue in seconds. I didn't allow her to make the actual changes, but she did have read access to everything, and helped me considerably.
So network-manager gets a new bug report about too-long non-ASCII AP names, I get online, and I don't know maybe Anthropic sneakily learned something from my local python projects. I am one of those vast-majority stupid people.
CamperBob2 14 hours ago [-]
This attitude makes zero sense to me. You benefit from the "training data" just like everyone else does. If you don't, that's a problem with you, not a problem with AI models.
AI solves exactly the meta-problem you describe: "I had a problem and needed to write a one-off doo-dad utility program to solve it." Now you can do something with your time besides writing pointless one-off doo-dads.
As for monetizing the training data, (a) it cost hundreds of millions of dollars to generate the weights, so why begrudge the companies that made the investment and did the research necessary to make it happen?; and (b) rest assured, whatever your doo-dad does, an open-weight model like GLM 5.2 can generate it for free using your own hardware.
So you don't have to pay anyone in that case. Well, except nVidia, I guess. Point granted there.
akudha 12 hours ago [-]
It does make sense. When podgietaru wrote their doo-dad programs, I guess they made their work available for humans to use and they didn't mind giving their work away for free. Some other person might say "I'm giving my work away free, I don't care if humans use it or AI, for any purpose". Both are legitimate choices, each according to their own.
As for "everyone benefiting from training data" - sure, but the AI companies are not investing Billions of dollars out of the goodness of their hearts, it is towards one and only singular goal of making profits (at some point). People might be sympathetic to these AI companies if they at least behave decently - they take everyone's work (text, software, fiction, music, images, videos...) without paying a penny to anyone. If they take everyone's work for free, they should give away anything that is built on that work also for free. This is before we even get to environment, privacy, hammering sites by not respecting robots.txt etc issues.
If I stole all veggies from your garden but spent time/money/effort making a meal, I should at minimum share it with you for free, no? Even if I spent my own money making the meal, it was made from stolen raw material...
protocolture 2 hours ago [-]
>As for "everyone benefiting from training data" - sure, but the AI companies are not investing Billions of dollars out of the goodness of their hearts, it is towards one and only singular goal of making profits (at some point).
Its not guaranteed they will ever get there, every day it seems increasingly likely that without massive government intervention we are just waiting for local open weights models to become popular.
Meanwhile my ISP gatekeeps the same free content behind a service payment. Their motive isnt free love and world peace, its also profit.
>If I stole all veggies from your garden but spent time/money/effort making a meal, I should at minimum share it with you for free, no?
If you cloned all the veggies in my garden you are free to clone them further and fill your belly you owe me nothing.
CamperBob2 10 hours ago [-]
If I stole all veggies from your garden but spent time/money/effort making a meal, I should at minimum share it with you for free, no?
Yes, and that's exactly my point. We are in violent agreement. It's shared with me, with you, with OP, and with everybody else. We can get a delicious meal for free or we can pay somebody else to serve us a slightly-tastier version.
Here, disregarding copyright law has fulfilled the very purpose of copyright law: to advance the useful arts and sciences. Copyright law was the best tool we had to accomplish that before, and now we have something better.
podgietaru 10 hours ago [-]
I wouldn’t be happy if those companies used my vegetables to open a Michelin star restaurant, even if I got a free McDouble out of the deal.
Plant based I guess I don’t know metaphors are hard.
CamperBob2 7 hours ago [-]
But the McDouble isn't an adequate comparison. Yes, the companies who opened Michelin-rated restaurants with the food they stole from your garden are giving you McDonalds'-level food for free and charging for the rest. Meanwhile, some other companies who raided your garden are giving you the plant-based equivalent of Ruth's Chris or Fogo de Chão for free.
The other thing is, after they stole all that stuff from your garden, it was somehow still there. Your neighbors on Hacker News say that some bandits raided your garden, but you can plainly see that no one has picked any fruit or uprooted any plants, and your security cameras reveal nothing more rapacious than a rabbit or two. You begin to suspect that your neighbors are gaslighting you.
foco_tubi 4 hours ago [-]
The plant metaphor breaks down when you pause, take a breath, put the McDouble down and realize that intellectual property is a completely different ownership concept than physical property.
cycomanic 10 hours ago [-]
But the AI companies are not publishing the models? Moreover they are charging for access to the models.
CamperBob2 10 hours ago [-]
Sure they are. Surf around on HuggingFace and you'll see dozens of open-weight models up to 120B parameters in size from for-profit US companies, freely downloadable (if not freely runnable, alas). Probably hundreds of them, at this point.
And a 3T parameter model is scheduled to be dropped by the Chinese on Monday.
They should certainly publish more, and if somebody were to argue that model weights trained by scraping copyrighted data should inherently be accessible to everyone, I'd be 100% in favor of that.
podgietaru 13 hours ago [-]
I don't know how else to explain it. It doesn't feel the same to have my contributions smushed into a linear-algebra machine? For me, the incentive was the idea that my contributions might have directly helped someone. That absolutely does not feel the same when I think "My answer has modified the back propagation of a Machine Learning training run."
It doesn't have to make sense to you - I just believe that I'm not exactly alone in this thought.
Now apply this to Art, free stories, writing etc. It feels bad to have your free contributions hoovered up and monetized. It doesn't feel particularly fair or ethical to me. And it'd make me double think before making something free and publicly available.
CamperBob2 13 hours ago [-]
Understood, there's certainly nothing invalid about your point of view here. It's just not one that I personally can come to terms with.
I've spent a lot of time in your shoes, wasting time on busy-work needed to accomplish a larger goal (and absolutely sharing the results freely, over multiple decades)... and I don't miss that part of it one bit.
jtuple 13 hours ago [-]
> Now you can do something with your time besides writing pointless one-off doo-dads
Writing software one-offs to scratch an itch was historically one of my most enjoyable past-times. AI trivializing that has been a very real theft of joy in my life.
Solving the actual problem was never the point, it was just motivation to do geek-out and craft some code.
AI is rapidly diminishing many interesting hobbies (coding, art, music, writing).
Having more free time when there's nothing fun nor exciting to do with it isn't really a benefit.
dminik 13 hours ago [-]
This is such a sad worldview. Life without any challenges. Every problem immediately solved without learning anything.
CamperBob2 10 hours ago [-]
What's sad is somebody giving you a godlike tool and watching you mope around muttering about "life without any challenges."
yuye 6 hours ago [-]
It's only considered "godlike" by those satisfied with mediocrity
CamperBob2 5 hours ago [-]
The person I replied to is apparently very content with mediocrity.
foco_tubi 8 hours ago [-]
The godlike tools are paid subscriptions at best and restricted to a handful of corpos by the government at best. The free scraps thrown to the proles are mostly novelty. The other day I asked 5.6 Sol to identify a push prop airplane and it said “this is a helicopter.”
yuye 6 hours ago [-]
>Now you can do something with your time besides writing pointless one-off doo-dads.
This is such a disgusting statement that goes against the very core of what open-source stands for. If even only one person got some use out of what they made, it is not pointless by definition.
mtVessel 13 hours ago [-]
| Now you can do something with your time besides writing pointless one-off doo-dads.
...thus eliminating all the tedious chatting, relaxing and making friends that people were previously forced to do while waiting for elevators.
CamperBob2 13 hours ago [-]
...thus eliminating all the tedious chatting, relaxing and making friends that people were previously forced to do while waiting for elevators
That also makes no sense, but I don't know what else I should have expected.
qotgalaxy 13 hours ago [-]
[dead]
primitivesuave 14 hours ago [-]
A couple years ago, I was running a website which allowed the public to view all the US government handouts to small businesses during the COVID-19 pandemic. It also tracked all the fraudulent loans being prosecuted by the DOJ, and allowed anyone to run structured queries over the public dataset. There was a "donate" button which took in ~$2k in donations over the lifetime of the site, and you could download the entire underlying dataset (around 10 GB uncompressed) for free directly on the site.
Despite the "download all data" link being prominently placed on the front page, the AI scrapers decided it would be more efficient to download terabytes upon terabytes of raw HTML by paginating through every possible facet on the search endpoint. Even with CloudFront caching results and a fairly efficient backend setup, the monthly bill ended up with around $1k just going toward network ingress/egress, so I shut down the site the following month.
deepsun 13 hours ago [-]
BigQuery has "public datasets", so users can even run complex SQL on it, but it's them who pays for it, not you. You only pay for data storage.
The crawlers would have still just hammered their site though, right?
primitivesuave 10 hours ago [-]
Yes, if I wanted to put a nice HTML interface over the query results then I would still end up with the same problem where some combinatoric explosion of query parameters to the `/search` endpoint, most of which are cache misses, leads to many many TBs of network egress.
pavel_lishin 12 hours ago [-]
I think the idea is that they could store the data in BigQuery, and point users of the site there.
dpoloncsak 12 hours ago [-]
Sure, but now you're moving the site from "Free data presented in a pleasant way to view" to a "pay-as-you-go database". Your audience shifts dramatically, and you lose the ability to share the data you're trying to present.
deepsun 5 hours ago [-]
No, it's free for everyone for the data sizes they have:
Free tier: 10 GB of active storage and 1 TiB of query data processed per month.
The crawlers would have still just hammered their site though, right?
40four 12 hours ago [-]
Why is your comment exactly word for word of another comment just one level above in the comment chain?
gorgonian 11 hours ago [-]
Maybe because they restated what they said instead of addressing to the previous commenter’s point.
40four 3 hours ago [-]
The comment I’m speaking of was from a different user, and it’s giving a strong smell that both users are bots unfortunately. Maybe I’m wrong, maybe it’s a coincidence those exact strings of characters were formed in isolation from each other in the same thread. I hope I’m wrong, but it’s a red flag
taneq 11 hours ago [-]
Because the crawlers would still have hammered their site, though.
(The GP post doesn’t actually meaningfully address the issue being raised. Adding BigQuery or whatever would not change the fact that (a) they already offered a method of getting all of the data in a cost effective way, and (b) the issue was that the crawlers hammered the site hard enough to make it economically unviable.)
lokar 12 hours ago [-]
I think snowflake has similar, you can rent it out or make it free
sillysaurusx 6 hours ago [-]
For what it’s worth, a Hetzner dedicated server has unlimited ingress and egress. It seems like it’s the only provider that does. Egress fees suck.
You can get a beefy one for about $40/mo on their server auction site.
Just... don’t miss payments. Ever. Or they’ll delete your server within a week or so.
userbinator 5 hours ago [-]
the AI scrapers decided it would be more efficient to download terabytes upon terabytes of raw HTML by paginating through every possible facet on the search endpoint.
No, they decided that would be a great way to convince you of the narrative and persuade you to pay for "security" services that further the incumbent browser monopoly.
They're not "AI scrapers", they're DDoS'ers manufacturing consent.
tailscaler2026 13 hours ago [-]
Building on AWS is a financial time bomb.
primitivesuave 10 hours ago [-]
Completely agree, I've worked at several companies where I saw the AWS bill balloon from five to six figures, usually because of pointless over-provisioning. However, it all became worth it when I saw Jeff Bezos go to space. [1]
Are you not able to put limits on how much the site can spend?
mystifyingpoi 11 hours ago [-]
As of 2026, still not, and probably never.
shermantanktop 12 hours ago [-]
Sure, but many a hobbyist has discovered the need for that the hard way.
13 hours ago [-]
BeeOnRope 13 hours ago [-]
Do you have any view on why the AI scrapers resulted in a heavier load than existing crawlers from eg search engines?
Where they more exhaustive or more frequent?
primitivesuave 10 hours ago [-]
They were both more exhaustive and more frequent. I don't remember the exact numbers, but it was definitely over 100x the traffic from search engines. By the way, search engines were allowed under robots.txt - I did want all 11.5 million loans to be individually indexed so they would pop up in Google search results, and I actually did receive/forward multiple tips about fraudulent loans because someone searched a business name on Google. All of this traffic was barely a blip, and my AWS bill for my hobby data science account was only ~$50-$100/month.
When I looked at the logs after getting the billing alert, 99.99% of the requests were to the "/search" endpoint with virtually every permutation of ~10-12 facets in the query parameters. There was only one scraper, but it triggered an enormous amount of network egress since it ended up missing the cache on the majority of queries.
pverheggen 12 hours ago [-]
There's been an explosion of vibe-coded scrapers that behave poorly and ignore robots.txt. Presumably OP disallowed crawling of the search endpoint for the reason stated (that crawling it would result in endless permutations of search filters.)
ratelimitsteve 13 hours ago [-]
not the OP but I'd say that google crawls you once and AIs scrape your page every time someone asks them a question that they think your page might be relevant to.
x3haloed 11 hours ago [-]
That’s a traffic design problem. You should be happy that your work is valuable and also protected it against excessive requests. Simple.
primitivesuave 9 hours ago [-]
I deployed a Lambda function behind CloudFront which rendered a simple HTML page with the query results from executing some SQL over the dataset. I served millions of page views for next to nothing because most pages were already in the cache.
I don't know where you get this expectation that people should anticipate that a crawler might try every possible combination of query parameters, thereby missing the cache on each one. Most people consider it a bitter and arrogant perspective, which is why this got downvoted.
breve 3 hours ago [-]
AI companies really do socialize the costs and privatize the profits.
Sites like The Numbers have to take on the cost of surviving the AI onslaught and the AI companies return nothing back to them.
Cthulhu_ 8 minutes ago [-]
To a point; as the post mentions, they added instructions for LLMs to their website and increased their licensing inquiries tenfold (the article didn't state anything about actual licensing payments being made though).
I don't think there would be any issues per se if the scrapers just paid licensing fees to get the good / complete data. But the issue was that they started to try and find exploits to get to data earlier.
jambalaya8 3 hours ago [-]
worse; it isn't just no return, it is also cost.
daniela-scott 1 hours ago [-]
[flagged]
djoldman 12 hours ago [-]
Just FYI, the bigger companies all allow you to block crawlers via robots.txt:
This only works for crawlers that play fair, unfortunately there's many that don't. Including ones that use consumer devices like TVs [0] and mobile phone apps that offer incentives to consumers (if they're even open about it) to use their internet connection to e.g. crawl websites. There's probably an army of hacked toothbrushes and the like doing the same thing.
I presume it would also cut you off even more from referral traffic.
pixelesque 9 hours ago [-]
Claude Bot still (at least last month, and it's been doing it for over a year now) seems to have a bug when traversing (at least my sites), wherein it drops the trailing slash of a directory (which is present in the a href tag), then makes the request to the subdirectory without the slash I put in the link, then Caddy automatically responds to that (via the built-in File handler) with a redirect telling it to add the trailing slash, and Claude Bot then makes the request again with the trailing slash that should have been there in the first place.
So most subdirectory URLs get two requests from Claude bot, the first one needless because that wasn't the URL in the tag.
globular-toast 54 minutes ago [-]
I won't eat your lunch if you put a sticker on your lunch box telling me not to.
spiderfarmer 11 hours ago [-]
90% of bot traffic on my network of websites is through headless Chrome, via residential bots nowadays. Impossible to block. Not even for Google, as they inflate my Adsense numbers as well.
6510 11 hours ago [-]
I once had the dumb idea to make websites in pdf but the idea is growing on me. Perhaps it should even be in animated gif with <img> <map>'s
ethagnawl 15 hours ago [-]
At the risk of oversimplifying things from a distance, this site -- especially the free, public-facing part of it -- seems like it would be an ideal candidate for a rewrite using static site generator/framework. That, coupled with a bot-aware CDN should keep them online at a reasonable cost for many years to come.
Otherwise, I'm very curious to know more about their old and new architecture and what sorts of mitigation/scaling strategies they've started using to keep the site online.
zackmorris 13 hours ago [-]
Ya this problem was solved 20 years ago with Varnish cache and Coral cache/CDN.
I think what's really going on is that bots expose how underpowered web servers has gotten in recent years. In the 2000s, even poorly-architected PHP sites tended to serve about 200 requests per second, with 1000+ being common for static sites. I remember when Node.js came out and claimed that it could serve more like 100,000 RPS due to its cooperative threading model. But today sites have a remarkable slowness to them, running many hundreds or thousands of database queries due to ORMs and N+1 problems, so that response times can be 500 ms or more and even 1000 simultaneous users stresses servers.
What went wrong is that nobody solved stuff like Russian doll caching in a general way separate from the programming language and database. We should have had ways to make dependency graphs using Etag headers as keys with real cache invalidation of dependent data. I went down that rabbit hole 10 years ago using touch events in Laravel with callbacks to handle cache invalidation when class model data was saved to the database. Also a query cache using Redis which I think might have been handled better at the database level anyway. After that experience, I can honestly say that cache invalidation is so difficult to get right that it's effectively an open problem. Meaning that programmings should use a package instead of rolling it by hand, and it should be a major concern from the start (along with sharding by user id or using something like Firebase).
Don't get me started on how the web should have been a P2P content-addressable memory anyway. Nearly everything should be available from a nearby edge peer, similarly to BitTorrent. But nobody bothered to solve how to make that work with HTTPS/SSL. I suspect that has to do with early flaws in the browser security model where the whole page has to be behind HTTPS or warnings appear. So it was never clear what was personally identifiable information (PII) or merely public data being served over HTTPS. To really solve that, we probably need real trust networks and maybe even zero-knowledge proofs.
Since these problems are so challenging to fix, and big companies can't be bothered to do it since they pulled the ladder up behind them, we're probably stuck with banal "are you human" challenge screens for the foreseeable future.
withinboredom 13 hours ago [-]
They are indeed quite challenging. We literally wrote a paper about it (https://arxiv.org/abs/2605.09114) ... you almost describe some of our original architecture too! You might find it interesting -- or not. https://getswytch.com
atherton94027 8 hours ago [-]
> I think what's really going on is that bots expose how underpowered web servers has gotten in recent years
More often than not they're running on a VPS, and cloud providers have been pushing the envelope of what a vCPU is. Amazon still bills hyperthreads as a single core!
Add to that their servers are often aging and you get a recipe for slow web services
Terr_ 11 hours ago [-]
> Nearly everything should be available from a nearby edge peer, similarly to BitTorrent.
Back in the days of Gnutella, I remember pushing people to use Magnet links [0] when sharing content.
> Ya this problem was solved 20 years ago with Varnish cache and Coral cache/CDN.
It was... if you are paying datacenter rates for the bandwidth
If you're paying cloud provider per GB pricing, nah, even text will add up if you happen to be targeted by a bunch of bots
> What went wrong is that nobody solved stuff like Russian doll caching in a general way separate from the programming language and database. We should have had ways to make dependency graphs using Etag headers as keys with real cache invalidation of dependent data.
we did that with nested ESI includes in Varnish so every "box" of content on the page was cached separately + some piping for invalidation, so if a given piece in the database was changed it sent invalidation to all nodes. There was also some grace so if the thing you wanted got updated RIGHT NOW you might get stale version while the new one is updated in background, and don't pay the latency cost
tengada1 14 hours ago [-]
Yeah this seems like one of the easiest access models to create an excellent security model for -- basically just static publishing. Sounds like they were in need of a rewrite anyway!
It probably seems daunting but to be honest this feels like a weekend's work at this point with LLM assistance. Not to be glib!
jaredwiener 14 hours ago [-]
But it also destroys the business model behind the site.
Sure, you could technically redesign to handle the bot traffic, but if the bot traffic is just taking the data and reducing any need for humans to visit the site, why is he putting in the effort to maintain the site?
tptacek 13 hours ago [-]
The business model of the site is apparently private data sales, not ad revenue.
jaredwiener 12 hours ago [-]
Bots don't make purchasing decisions.
Terr_ 11 hours ago [-]
Some people are already trying to make that happen, but I'm convinced it'll be a net-negative.
jaredwiener 11 hours ago [-]
Rephrase: Bots SHOULDNT make purchasing decisions.
fragmede 5 hours ago [-]
Because... ? If I'm operating a site, and I want $X to allow a bot to scrape my site, why shouldn't the bot be allowed to make the purchasing decision to scrape the site? Obviously the bot owner would have their own set of guardrails, but if it allows sites like The numbers to stay up because bots aren't going to look at ads so the old model of displaying ads isn't going to work, I have a hard time seeing that as a bad thing.
jaredwiener 4 hours ago [-]
I mean it from the consumer perspective.
hyperhello 14 hours ago [-]
Well, why is he? If it's for ad revenue, this problem is surely global to the web. If it's because he likes to do it, it shouldn't matter if one AI or a trillion scrape the site. I feel like we badly need a rethink of the web architecture after DoubleClick anyway. Maybe the name of the game should be to cut down a site's assets very hard and use static hosting for them. This interferes with crummy sites that show a mess of inlined ads every refresh, but now that there are a billion 'poor users' this is no longer feasible.
jaredwiener 14 hours ago [-]
Maybe more an inherent problem with these chatbots.
LLMs are great, but they aren't producing new information. You still need people for that. But if you cut down any incentive for the people to do that, the LLMs will starve.
RobotToaster 14 hours ago [-]
Or just use a cache.
theragra 10 hours ago [-]
I have small website with archive of older radio broadcasts mostly in Russian. Lately, 90% of traffic comes from USA ;)
Thanks God i can serve up to a terabyte per month of traffic easily, otherwise it would be a catastrophe. I am not against bot scraping, but I worry about stability for meat visitors.
So, I think of enabling payments for website visits cloudflare recently developed.
I also have the problem of old technology like the mentioned site. While my personal blog uses static generator, archive website uses ancient Drupal version, which has no security patches for many years already.
gajus 14 hours ago [-]
What a throwback. Back in 2015 I have started Applaudience, which at the time was the only provider of real-time cinema ticket sales data. I still remember comparing our numbers against TheNumbers.com as part of calibration. I have since moved on to other businesses, but this remains one of my favorite pieces of technology that I have developed. Would love to bring it back one day.
rodarmor 5 hours ago [-]
There’s an easy solution: Bruce should place large bets on the prediction markets right before publishing the relevant data. It is legal, ethical, and, for him, risk free.
ajkjk 12 hours ago [-]
It feels like there is fundamentally missing infrastructure here that is needed to make these problems go away.
Basically bots need to be (somehow) paying for the traffic they create, or prevented from creating it, or told to go away and then fined if they violate the request. No idea how to do these or even at what level in the stack they should happen, but they need to happen eventually somehow.
polio 12 hours ago [-]
Every website I visit could get a fraction of a cent in my Cloudflare Wallet. A human browsing incurs a few dollars a month. Plus, any website I access in this way decides not to show me ads either and just charges me the cost of serving the data plus a nominal profit margin.
Disclosure: I am a shareholder and would love for them to solve the AI bot problem and the ad problem like this.
wredcoll 8 hours ago [-]
Maybe we'll all eat our words and cryptocurrencies will actually become useful for something.
weird-eye-issue 7 hours ago [-]
This does not require cryptocurrency
john_strinlai 12 hours ago [-]
we've been working on basically the same problem in the email space for a few decades (legit email vs. mass spam). its a very hard (i think impossible) problem.
ajkjk 4 hours ago [-]
Well, it is technologically not that hard to imagine a solution---the problem is social, getting everyone to agree on how to do it, figure out the policy around it, etc. And the situation is getting so bad that maybe it is time for someone (maybe someone reading this thread) to figure it out.
4 hours ago [-]
yuye 6 hours ago [-]
The same goes for telephony.
I reject all phone calls by default, unless I'm expecting a call.
paxys 14 hours ago [-]
I know people have opinions about Cloudflare but why not use it here, at least as a stop gap? Stopping bot traffic is one thing it does very well.
philipkglass 9 hours ago [-]
I set Cloudflare up a couple of months ago specifically to block bot traffic. It didn't do anything for me. Dumb bots were still hitting every special link on my wiki fast enough that the server was continually swamped running Lua scripts. 65% of the traffic for my English-language site was coming from Vietnam. But I didn't want to block Vietnam altogether, because my hobby site has genuine users from there too.
I eventually just shut down my mediawiki instance. I couldn't find a way to keep it online and still run on an affordable VPS.
spiderfarmer 11 hours ago [-]
It tried it. Hundreds of “genuine” visitors per day on a new website with no search engine presence and no links. That’s a very leaky fence..
paxys 10 hours ago [-]
Hundreds is better than billions.
NetMageSCW 13 hours ago [-]
Cost?
esseph 13 hours ago [-]
It's free for this
cogogo 12 hours ago [-]
I think I probably use AI like a lot of consumers out there. Search engines have gotten bad and AI really good at answering fairly specific questions. Often pointing at sites like wikipedia. Pretty clear changing my behavior will have zero impact but it is certainly part of the problem. Feels a lot like my CO2 consumption.
*Edit - CO2 creation
spiderfarmer 11 hours ago [-]
GPTBot crawled millions of pages on my website. I get a handful of visitors from them. Meanwhile Google sends me 10k visitors each day.
So GPTBot is now blocked.
Centigonal 4 hours ago [-]
Our internet ecosystem is becoming more and more hostile to open information commons. I like open information commons -- what do we do about this?
jagermo 47 minutes ago [-]
prediction markets were a huge mistake.
Drdiamond 4 hours ago [-]
This is such a sad worldview. Life without any challenges and solved using AI. it will vastly decrease our learning
11 hours ago [-]
datadrivenangel 14 hours ago [-]
If the AI companies destroy the open web, eventually they'll need to start curating knowledge sources just like netflix makes movies and amazon has physical stores...
nradov 14 hours ago [-]
Already happening. Frontier LLM vendors have been hiring human domain experts specifically to create their own proprietary training data in targeted verticals.
xeyownt 11 hours ago [-]
Prediction market as well as anything related to gambling must disappear from earth surface. Full stop.
tehjoker 14 hours ago [-]
The real story here is that prediction markets were banned for a reason and loosening the rules is causing chaos just as was expected. AI plays little role in this story, hacking by humans would also be motivated by financial returns, unless the element is that AI hacking is cheaper and the returns are not so big.
tailrecursion 11 hours ago [-]
Do the AI labs in the U.S. publish the AWS IP address sets that their crawlers use? Is that not an effective way to block those crawlers anymore?
theragra 10 hours ago [-]
They are often from residential IP.there are even businesses that allow you renting such IPs
weird-eye-issue 7 hours ago [-]
Nobody serious would ever use an AWS IP lol
hebleb 14 hours ago [-]
Good read, I was so curious how this happened a few months ago
advisedwang 13 hours ago [-]
This article, and possibly the owners of the site, seem to muddle together several problems:
1. bot traffic causing infrastructure cost
2. scraping circumventing paying for licenses
3. the risk of hacking
shermantanktop 12 hours ago [-]
You missed the speculated motivation: unrestricted prediction markets, which are an open and broad incentive to do whatever actions might provide a slight edge in betting.
Webscraping is among the more benign things that gambling-on-anything can drive. And even that has a negative impact, as seen here.
I have no idea why those sites are legal.
dumberquestions 15 hours ago [-]
Ironic that polymarkets were being advertised as helping society make better predictions.
jambalaya8 14 hours ago [-]
John Brunner and Alvin Toffler both made stark warnings wrapped in futuristic giddiness about things like it. People like Fuller no doubt thought polling on a large scale was terrific. There were old usenet groups and BBS subs (minus the money aspect) experimenting with the model. I do not believe they ever are or were good in a largescale model (money or not).
mrandish 13 hours ago [-]
I mean wisdom of crowds, super-forecasters, calibration and pre-registration are useful tools that can result in better predictions, turning it into online gambling is where it went sideways.
Terr_ 11 hours ago [-]
Mini-rant: Promoters claim that the system serves a public good by helping society discover/converge on useful truths sooner.
Even where the question/event is of public interest, the opposite happens instead. People with expertise or non-public information are incentivized to misdirect and delay as long as possible, as that maximizes what they can make from betting.
hyperhello 14 hours ago [-]
If the scrapers are going to get it anyway, put the data up as a zip somewhere.
jjgreen 14 hours ago [-]
It would not make any difference.
codemonkey-zeta 14 hours ago [-]
Indeed, the article mentions Wikipedia experiencing similar scraping pains, even though they already DO have bulk data available.
HeatrayEnjoyer 13 hours ago [-]
Who are running these bots? I presume developers at all of the frontier labs know (or at least would know to look for) Wikipedia has bulk APIs for automated access. Unnecessary scraping increases their workload/costs too, so why in 2026 is this still a problem?
esseph 13 hours ago [-]
Black market and gray market data. All the firms want data. All the other firms want data. The banks want data. The other criminals also want data for their crimes and schemes. Oh insurance companies, and the ATS systems. Everybody wants as much data as they can get and they don't care how they get it.
esseph 11 hours ago [-]
This data selling also happens with leaks of all kinds like medical data, often to current or future employers, health insurance companies, etc.
antisthenes 12 hours ago [-]
> Read the Docs, a non-profit that hosts documentation for open-source software, who watched a single crawler download 73 terabytes of zipped HTML in one month, costing it over $5,000 in bandwidth
From the article.
Not the same site, but an example of the same issue.
puskavi 6 hours ago [-]
Eh, wouldnt just adding some POW challenge, like anubis solve the scraping by making it hurt crawlers wallets?
14 hours ago [-]
freediddy 11 hours ago [-]
All new content will be behind paywalls. This will be the new way the Internet works and I've said this for the last several years. There's no value in writing unique content just to have AI steal it and distribute it without you getting any clicks. The only way it works if it there's a licensing agreement and if they don't want to pay, then who cares, you weren't going to make any money from them anyway.
npilk 11 hours ago [-]
What if you aren't motivated solely by making money? There will still be new free content created by hobbyists, passionate creators, altruists, etc.
(Agree with you more generally.)
freediddy 11 hours ago [-]
Why would you produce content and have no one read it or visits your site, but OpenAI and Anthropic make millions from it? At some point it becomes stupid to just give free money to these companies when they steal literally all your traffic and content. As per the article, Anthropic sends 1 page view for every 38,000 views they get.
ssl-3 10 hours ago [-]
Very few people showed up to read stuff on the old web, and yet: It existed.
jambalaya8 14 hours ago [-]
I always liked this site, but reading this and seeing the anger about expecting the site maintainer to do things for you is repulsive. Frankly, if he wanted to pull his site down with no notice that is perfectly within his right. It was/is his site. He doesn't owe anyone a .tar.gz either. His work.
BraveOPotato 14 hours ago [-]
I agree. I've seen it happen before on a project I use. I decided to take a look at the repo for one of the plugins, and I saw a heinous issue that basically was TELLING (not even asking) the maintainer to fix it.
Deplorable behavior indeed
jambalaya8 14 hours ago [-]
Like dominoes, as soon as it is accepted in a few places, people think it is acceptable to push to 'share'. It's almost terroristic sometimes, the pressure some maintainers are under.
ropable 8 hours ago [-]
This is another reason we can't have nice things. The level of entitlement required to send an angry message to someone complaining about their free resource being offline is mind-boggling, though.
BrenBarn 4 hours ago [-]
Every example of this confirms my view that we should treat the spread of these AI scrapers like we treat the proliferation of drugs. We should be seeking to bust AI rings like we seek to bust drug cartels.
NetMageSCW 13 hours ago [-]
What does a cyber attack have to do with AI scraping?
john_strinlai 13 hours ago [-]
they both have a risk of harm which the site operator was no longer comfortable with.
brcmthrowaway 15 hours ago [-]
Its just too easy for a technically minded bored person to produce slop that hammers websites. There needs to be a penalty for this.
cawksuwcka 7 hours ago [-]
[dead]
draw_down 15 hours ago [-]
Sorry! Just yesterday many of us decided that AI scraping isn't a real problem, and anytime it is blamed, it's a cover for something else.
They should sue Anthropic for this distillation attack
vachina 13 hours ago [-]
> The world we have built thus far is so incredibly ill-prepared for the power and scale of the AI models we all have access to.
Exactly, so use the AI to secure your servers. Ask the AI to audit your site for any security holes. If unable to rewrite, at least harden the existing code. AI’s are really really cheap (and fast) security consultants now.
nater5000 14 hours ago [-]
This is an odd example to present this argument through.
I'm not defending AI bots overwhelming websites or hackers motivated by Polymarket, etc., but I don't really believe a 30 year old website with "approximately 160,000 source files serving around 2 million pages" is a good litmus test for the state of the online world. What's worse is the absurdity that this basic site offering niche data would be targeted because Polymarket depends on it for some of their bets, something that most websites don't have to deal with. Frankly, that seems like a much more interesting angle to explore than "this old website that should have been re-written multiple times over the last three decades now doesn't have a choice but to be re-written."
Oddly, it seems like the solution, at least in this specific case, is relatively simple (and ironic): pay for an LLM to re-write this site in a modern, scalable, secure way and have that LLM monitor the site to ensure it is behaving correctly. Put this thing in a modern platform behind a proper cache (i.e., throw the whole thing up in Cloudflare) and many of these problems just don't exist anymore. If this site is as basic as it seems, this could be a weekend project.
It's certainly an interesting story, and I don't blame the owner of The Numbers for handling his site the way he has, but framing this as "AI is destroying our beloved internet" just seems obtuse, at least through this lens.
jambalaya8 14 hours ago [-]
The actual solution is damn simple and painfully obvious: make things like so-called "prediction markets" illegal.
The intelligence community came to the conclusion (after research and experiments) that things like "prediction markets" were a bad idea in 1996.
Far worse now than then, with the net and AI of 2026 and the insane number of people now online that want to make an easy buck.
Don't get me wrong; I am guessing some of us on here would do well on those places. But they should not exist.
14 hours ago [-]
jaredwiener 14 hours ago [-]
Isn't this the NRA's argument? The only way to stop a bad guy with a gun is a good guy with a gun.
Or at least somewhere between that and full on protection racket.
LLMs come on the scene, hammer the site until it goes offline or racks up bills that threaten bankruptcy, and the solution then is to pay for an LLM to fix it and monitor it.
That's a real nice site ya got there, it would be a shame if massive datacenters going up around the world were to start hammering it from tens of thousands of IP addresses....
Millions of small sites and creators don’t have the ability to design their own protections against large scale automated access. If useful content now attracts aggressive crawler traffic, many sites will be too expensive or unreliable to run.
Is part of the answer a community response? Perhaps a community-maintained toolkit, based on traffic data, that host sites can apply? It could include standard agent identification, rate limiting, traffic classification, access policies, caching, challenge mechanisms, logging, attribution and usage control etc etc.
In effect, we need much stronger road rules for today’s automated traffic, available as open technical patterns and libraries rather than every site owner having to invent this alone (they won’t).
The "good bots" which identify themselves and follow the rules are easy enough to block, so not a problem. It is the "bad bots" which pretend to be real users and hide behind residential proxies, and so are almost impossible to block at the moment, which are the problem.
Given that there are big companies openly (i.e. on the clearweb, not even darkweb) selling access to these residential proxy botnets of compromised smart TVs[0] and mobile phones and other devices, can we not simply get a database of the these IPs and block access from them? I would venture that almost every single one of those residential users are unaware that they have devices in their homes which have been compromised and are being abused in this way, so if they were to start seeing messages from more and more sites along the lines of "Access to this site has been blocked because unusual traffic has been detected from your computer network. Please check all devices on your network and remove any malware which may be routing this traffic." then maybe we could start addressing the problem at the source.
[0] https://news.ycombinator.com/item?id=49000864
Which brings us to the old fashioned flood control mechanisms. That is, the toolkit you propose already exists and has been used for decades in various iterations to protect against various forms of attack (slashdotting, DDOS attacks, overzealous search engines, and now AI scrapers).
Have you looked into those before? Companies like Cloudflare have been at the forefront of this field for a long time now.
The article says (speculates?) that malicious users are trying to get privileged access for an edge in prediction market betting. From the article:
> If you could see The Numbers data before everyone else, every single week, you would have a significant edge over all the other traders - learning the answers slightly ahead of publication would allow you to front-run the trades.
You can unilaterally stop trading 'before a market closes with some time buffer to prevent this.' No need for centralised action.
See also stock market trading, where companies would go to great lengths to shave nanoseconds off of getting information and putting in trades. Now apply this min/maxing to worldwide, unregulated and anonymous.
This is what tech libertarians / cryptobros want.
I know this isn't really the point of the article, but I've been thinking about this a lot. I wonder if we're going to see more resources go this way.
I used to publish little doo-dads as opensource software. Not because it was something that was legitimately ground breaking or anything (it absolutely wasn't) but because I had a problem, and I thought "heh wouldn't it be cool if someone else had a similar problem and could use my resource for it."
But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.
And I can see why maintainers of sites like this, or other free but incredibly useful resources might start to get irate at that.
The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem.
> And I can see why maintainers of sites like this, or other free but incredibly useful resources might start to get irate at that.
Those who object to the scraping fall into several camps, but the biggest complaint I am hearing is that it increases both maintenance costs and time. In other words: it sucks when people are using your work in a manner that you find offensive, but it goes beyond that by doing actual harm.
> In other words: it sucks when people are using your work in a manner that you find offensive
That to me still reeks of Dog in the Manger mentality. If you publish something for the world to use, you should neither care nor even track, much less discriminate by (or suddenly seek compensation for) who is using it.
But people are free to not publish things or post things online with a more restrictive license. Not that a license stops things from being indexed.
People doing it with a couple of machines and residential proxies versus Anthropic doing it with 2 data centers worth of machines.
Scale matters.
I used to rank pretty well trashing crappy credit cards and encouraging people to switch to better options, then Google decided that 10 results for the card issuers website was better.
Same shit when I manually wrote proto-gethuman posts on calling telecoms/banks/etc (and also pushed visitors to try an Indy ISP or credit unions), then Google felt it was better to drive users to the telecom’s website that wants you to do anything but call them.
Please do train on my pre-LLM gold!
(I never publish my recommendations because they seem to complicated for people. Like using 30 GB plan/$10/monthly from T-Mobile for data and then using Tello for $10 plan for voice/text. This would require a 2 esim/sim card phone. I am currently using a Moto G Power 2024 phone from eBay $90/new, which was better than the $200 slightly used Pixel 6a from Swappa. Most people would just save the hassle, get a Galaxy phone with a phone contract.)
If you’re on page 2 of the results, you effectively don’t exist so I stopped bothering/benefiting from display ads.
Coincidentally, I did try to get deeper into the cellular service side (it’s another high margin and high customer value segment), but I did better on the finance side.
I suspect the next AGPL (if not normal GPL) will explicitly prohibit training of closed models against the source material. Ideally it will prohibit training of open weight models too. Either share the entire process, or go piss up a rope.
I dont get this.
You provided something for free to help people, but dont want to do that anymore because it might go into training data and help many many more people?
So far LLMs have been loan funded donations of loss leading services. They might never actually make their first dollar of profit.
Meanwhile ISPs the world over have been monetizing access to your content.
I feel like what you mean is that you want control over attribution.
I feel the same way too. But guess what, all the code you did not publish gets into the training corpus anyway (when you gave Codex or whatever full read access to your filesystem).
I recently had trouble connecting to a new AP on my laptop. After a quarter hour of frustration, I connected to the AP of my phone, asked Claude Code what the problem is, and she found the issue in seconds. I didn't allow her to make the actual changes, but she did have read access to everything, and helped me considerably.
So network-manager gets a new bug report about too-long non-ASCII AP names, I get online, and I don't know maybe Anthropic sneakily learned something from my local python projects. I am one of those vast-majority stupid people.
AI solves exactly the meta-problem you describe: "I had a problem and needed to write a one-off doo-dad utility program to solve it." Now you can do something with your time besides writing pointless one-off doo-dads.
As for monetizing the training data, (a) it cost hundreds of millions of dollars to generate the weights, so why begrudge the companies that made the investment and did the research necessary to make it happen?; and (b) rest assured, whatever your doo-dad does, an open-weight model like GLM 5.2 can generate it for free using your own hardware.
So you don't have to pay anyone in that case. Well, except nVidia, I guess. Point granted there.
As for "everyone benefiting from training data" - sure, but the AI companies are not investing Billions of dollars out of the goodness of their hearts, it is towards one and only singular goal of making profits (at some point). People might be sympathetic to these AI companies if they at least behave decently - they take everyone's work (text, software, fiction, music, images, videos...) without paying a penny to anyone. If they take everyone's work for free, they should give away anything that is built on that work also for free. This is before we even get to environment, privacy, hammering sites by not respecting robots.txt etc issues.
If I stole all veggies from your garden but spent time/money/effort making a meal, I should at minimum share it with you for free, no? Even if I spent my own money making the meal, it was made from stolen raw material...
Its not guaranteed they will ever get there, every day it seems increasingly likely that without massive government intervention we are just waiting for local open weights models to become popular.
Meanwhile my ISP gatekeeps the same free content behind a service payment. Their motive isnt free love and world peace, its also profit.
>If I stole all veggies from your garden but spent time/money/effort making a meal, I should at minimum share it with you for free, no?
If you cloned all the veggies in my garden you are free to clone them further and fill your belly you owe me nothing.
Yes, and that's exactly my point. We are in violent agreement. It's shared with me, with you, with OP, and with everybody else. We can get a delicious meal for free or we can pay somebody else to serve us a slightly-tastier version.
Here, disregarding copyright law has fulfilled the very purpose of copyright law: to advance the useful arts and sciences. Copyright law was the best tool we had to accomplish that before, and now we have something better.
Plant based I guess I don’t know metaphors are hard.
The other thing is, after they stole all that stuff from your garden, it was somehow still there. Your neighbors on Hacker News say that some bandits raided your garden, but you can plainly see that no one has picked any fruit or uprooted any plants, and your security cameras reveal nothing more rapacious than a rabbit or two. You begin to suspect that your neighbors are gaslighting you.
And a 3T parameter model is scheduled to be dropped by the Chinese on Monday.
They should certainly publish more, and if somebody were to argue that model weights trained by scraping copyrighted data should inherently be accessible to everyone, I'd be 100% in favor of that.
It doesn't have to make sense to you - I just believe that I'm not exactly alone in this thought.
Now apply this to Art, free stories, writing etc. It feels bad to have your free contributions hoovered up and monetized. It doesn't feel particularly fair or ethical to me. And it'd make me double think before making something free and publicly available.
I've spent a lot of time in your shoes, wasting time on busy-work needed to accomplish a larger goal (and absolutely sharing the results freely, over multiple decades)... and I don't miss that part of it one bit.
Writing software one-offs to scratch an itch was historically one of my most enjoyable past-times. AI trivializing that has been a very real theft of joy in my life.
Solving the actual problem was never the point, it was just motivation to do geek-out and craft some code.
AI is rapidly diminishing many interesting hobbies (coding, art, music, writing).
Having more free time when there's nothing fun nor exciting to do with it isn't really a benefit.
This is such a disgusting statement that goes against the very core of what open-source stands for. If even only one person got some use out of what they made, it is not pointless by definition.
...thus eliminating all the tedious chatting, relaxing and making friends that people were previously forced to do while waiting for elevators.
That also makes no sense, but I don't know what else I should have expected.
Despite the "download all data" link being prominently placed on the front page, the AI scrapers decided it would be more efficient to download terabytes upon terabytes of raw HTML by paginating through every possible facet on the search endpoint. Even with CloudFront caching results and a fairly efficient backend setup, the monthly bill ended up with around $1k just going toward network ingress/egress, so I shut down the site the following month.
Free tier: 10 GB of active storage and 1 TiB of query data processed per month.
https://github.com/phiresky/sql.js-httpvfs
https://duckdb.org/2021/10/29/duckdb-wasm
(The GP post doesn’t actually meaningfully address the issue being raised. Adding BigQuery or whatever would not change the fact that (a) they already offered a method of getting all of the data in a cost effective way, and (b) the issue was that the crawlers hammered the site hard enough to make it economically unviable.)
You can get a beefy one for about $40/mo on their server auction site.
Just... don’t miss payments. Ever. Or they’ll delete your server within a week or so.
No, they decided that would be a great way to convince you of the narrative and persuade you to pay for "security" services that further the incumbent browser monopoly.
They're not "AI scrapers", they're DDoS'ers manufacturing consent.
1. https://www.youtube.com/watch?v=IOmX793-5t4
Where they more exhaustive or more frequent?
When I looked at the logs after getting the billing alert, 99.99% of the requests were to the "/search" endpoint with virtually every permutation of ~10-12 facets in the query parameters. There was only one scraper, but it triggered an enormous amount of network egress since it ended up missing the cache on the majority of queries.
I don't know where you get this expectation that people should anticipate that a crawler might try every possible combination of query parameters, thereby missing the cache on each one. Most people consider it a bitter and arrogant perspective, which is why this got downvoted.
Sites like The Numbers have to take on the cost of surviving the AI onslaught and the AI companies return nothing back to them.
I don't think there would be any issues per se if the scrapers just paid licensing fees to get the good / complete data. But the issue was that they started to try and find exploits to get to data earlier.
[0] https://news.ycombinator.com/item?id=49000864
I presume it would also cut you off even more from referral traffic.
So most subdirectory URLs get two requests from Claude bot, the first one needless because that wasn't the URL in the tag.
Otherwise, I'm very curious to know more about their old and new architecture and what sorts of mitigation/scaling strategies they've started using to keep the site online.
I think what's really going on is that bots expose how underpowered web servers has gotten in recent years. In the 2000s, even poorly-architected PHP sites tended to serve about 200 requests per second, with 1000+ being common for static sites. I remember when Node.js came out and claimed that it could serve more like 100,000 RPS due to its cooperative threading model. But today sites have a remarkable slowness to them, running many hundreds or thousands of database queries due to ORMs and N+1 problems, so that response times can be 500 ms or more and even 1000 simultaneous users stresses servers.
What went wrong is that nobody solved stuff like Russian doll caching in a general way separate from the programming language and database. We should have had ways to make dependency graphs using Etag headers as keys with real cache invalidation of dependent data. I went down that rabbit hole 10 years ago using touch events in Laravel with callbacks to handle cache invalidation when class model data was saved to the database. Also a query cache using Redis which I think might have been handled better at the database level anyway. After that experience, I can honestly say that cache invalidation is so difficult to get right that it's effectively an open problem. Meaning that programmings should use a package instead of rolling it by hand, and it should be a major concern from the start (along with sharding by user id or using something like Firebase).
Don't get me started on how the web should have been a P2P content-addressable memory anyway. Nearly everything should be available from a nearby edge peer, similarly to BitTorrent. But nobody bothered to solve how to make that work with HTTPS/SSL. I suspect that has to do with early flaws in the browser security model where the whole page has to be behind HTTPS or warnings appear. So it was never clear what was personally identifiable information (PII) or merely public data being served over HTTPS. To really solve that, we probably need real trust networks and maybe even zero-knowledge proofs.
Since these problems are so challenging to fix, and big companies can't be bothered to do it since they pulled the ladder up behind them, we're probably stuck with banal "are you human" challenge screens for the foreseeable future.
More often than not they're running on a VPS, and cloud providers have been pushing the envelope of what a vCPU is. Amazon still bills hyperthreads as a single core!
Add to that their servers are often aging and you get a recipe for slow web services
Back in the days of Gnutella, I remember pushing people to use Magnet links [0] when sharing content.
[0] https://en.wikipedia.org/wiki/Magnet_URI_scheme
It was... if you are paying datacenter rates for the bandwidth
If you're paying cloud provider per GB pricing, nah, even text will add up if you happen to be targeted by a bunch of bots
> What went wrong is that nobody solved stuff like Russian doll caching in a general way separate from the programming language and database. We should have had ways to make dependency graphs using Etag headers as keys with real cache invalidation of dependent data.
we did that with nested ESI includes in Varnish so every "box" of content on the page was cached separately + some piping for invalidation, so if a given piece in the database was changed it sent invalidation to all nodes. There was also some grace so if the thing you wanted got updated RIGHT NOW you might get stale version while the new one is updated in background, and don't pay the latency cost
It probably seems daunting but to be honest this feels like a weekend's work at this point with LLM assistance. Not to be glib!
Sure, you could technically redesign to handle the bot traffic, but if the bot traffic is just taking the data and reducing any need for humans to visit the site, why is he putting in the effort to maintain the site?
LLMs are great, but they aren't producing new information. You still need people for that. But if you cut down any incentive for the people to do that, the LLMs will starve.
Previously it was like 5%.
URL is http://radar.lv btw
Thanks God i can serve up to a terabyte per month of traffic easily, otherwise it would be a catastrophe. I am not against bot scraping, but I worry about stability for meat visitors.
So, I think of enabling payments for website visits cloudflare recently developed.
I also have the problem of old technology like the mentioned site. While my personal blog uses static generator, archive website uses ancient Drupal version, which has no security patches for many years already.
Basically bots need to be (somehow) paying for the traffic they create, or prevented from creating it, or told to go away and then fined if they violate the request. No idea how to do these or even at what level in the stack they should happen, but they need to happen eventually somehow.
Disclosure: I am a shareholder and would love for them to solve the AI bot problem and the ad problem like this.
I reject all phone calls by default, unless I'm expecting a call.
I eventually just shut down my mediawiki instance. I couldn't find a way to keep it online and still run on an affordable VPS.
*Edit - CO2 creation
So GPTBot is now blocked.
1. bot traffic causing infrastructure cost
2. scraping circumventing paying for licenses
3. the risk of hacking
Webscraping is among the more benign things that gambling-on-anything can drive. And even that has a negative impact, as seen here.
I have no idea why those sites are legal.
Even where the question/event is of public interest, the opposite happens instead. People with expertise or non-public information are incentivized to misdirect and delay as long as possible, as that maximizes what they can make from betting.
From the article.
Not the same site, but an example of the same issue.
(Agree with you more generally.)
Deplorable behavior indeed
https://news.ycombinator.com/item?id=49005747
Exactly, so use the AI to secure your servers. Ask the AI to audit your site for any security holes. If unable to rewrite, at least harden the existing code. AI’s are really really cheap (and fast) security consultants now.
I'm not defending AI bots overwhelming websites or hackers motivated by Polymarket, etc., but I don't really believe a 30 year old website with "approximately 160,000 source files serving around 2 million pages" is a good litmus test for the state of the online world. What's worse is the absurdity that this basic site offering niche data would be targeted because Polymarket depends on it for some of their bets, something that most websites don't have to deal with. Frankly, that seems like a much more interesting angle to explore than "this old website that should have been re-written multiple times over the last three decades now doesn't have a choice but to be re-written."
Oddly, it seems like the solution, at least in this specific case, is relatively simple (and ironic): pay for an LLM to re-write this site in a modern, scalable, secure way and have that LLM monitor the site to ensure it is behaving correctly. Put this thing in a modern platform behind a proper cache (i.e., throw the whole thing up in Cloudflare) and many of these problems just don't exist anymore. If this site is as basic as it seems, this could be a weekend project.
It's certainly an interesting story, and I don't blame the owner of The Numbers for handling his site the way he has, but framing this as "AI is destroying our beloved internet" just seems obtuse, at least through this lens.
The intelligence community came to the conclusion (after research and experiments) that things like "prediction markets" were a bad idea in 1996.
Far worse now than then, with the net and AI of 2026 and the insane number of people now online that want to make an easy buck.
Don't get me wrong; I am guessing some of us on here would do well on those places. But they should not exist.
Or at least somewhere between that and full on protection racket.
LLMs come on the scene, hammer the site until it goes offline or racks up bills that threaten bankruptcy, and the solution then is to pay for an LLM to fix it and monitor it.
That's a real nice site ya got there, it would be a shame if massive datacenters going up around the world were to start hammering it from tens of thousands of IP addresses....