SShortSingh.
Back to feed

15 AI Crawlers Explained: What Blocking Each One Actually Costs You

0
·1 views

A verified registry of 15 AI bots has been published, categorizing them into training crawlers, assistant crawlers, search crawlers, and scrapers — with each category carrying different consequences for blocking. A key correction highlighted is that Google-Extended and Applebot-Extended are not crawlers but robots.txt tokens that control AI training data use, not search indexing or AI Overviews. Blocking assistant crawlers, which fetch pages to answer live user queries, removes a site from AI-generated answers, while blocking training crawlers only affects model training with no traffic impact. Cloudflare introduced a dashboard toggle in mid-2025 to manage AI crawler rules without manual robots.txt editing, though the underlying token logic remains unchanged. The guide urges website owners to make bot-by-bot blocking decisions rather than blanket blocks, which sacrifice AI-answer visibility without necessarily achieving the intended opt-outs.

Read the full story at DEV Community

This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)

Log in to join the discussion and vote.

Log in

Related stories

0
ProgrammingHacker News ·

PostHog Newsletter Explores Limits of Delegating Tasks to AI Agents

PostHog published a newsletter piece examining how much work can realistically be handed off to autonomous AI agents. The article explores the boundaries of agent autonomy in software and product development workflows. It was shared on Hacker News, where it received 3 points at the time of reporting. The discussion thread had no comments, suggesting the post was relatively new or niche in appeal.

0
ProgrammingDEV Community ·

How to Deploy, Update, and Scale a Custom App on an On-Premise Kubernetes Cluster

Part 6 of this Kubernetes on-premise series walks through building a Docker image for a custom Spring Boot REST API and publishing it to Docker Hub or a private registry. The guide demonstrates creating a Deployment using the kubectl command line and exposing it via a LoadBalancer Service with a manually assigned external IP, as on-premise environments lack cloud-native load balancer provisioning. It covers rolling updates using kubectl set image, which replaces the running container image with a new version while recording the change in revision history. In case of issues, a rollback to the previous revision can be performed instantly with kubectl rollout undo. The article also explains both manual replica scaling with kubectl scale and automatic scaling through a Horizontal Pod Autoscaler configured to maintain CPU utilization below a defined threshold.

0
ProgrammingDEV Community ·

On-Premise Kubernetes Cluster: Deploying Your First Container with Nginx

A technical tutorial series on building an on-premise Kubernetes cluster has reached its fifth installment, focusing on deploying the first application. Using Nginx as a practical example, the guide walks through creating a Deployment manifest with two pod replicas managed across available worker nodes. A NodePort Service is then configured to expose the application externally via any cluster node's IP address, without requiring an external load balancer. The article also covers best practices such as organizing Kubernetes manifests in per-application directories for easier maintenance and version control. Successfully accessing the default Nginx page through the assigned NodePort confirms the cluster is functioning end-to-end.

0
ProgrammingDEV Community ·

Guide: Installing Containerd and Kubernetes on an On-Premise Cluster

A technical tutorial series on building an on-premise Kubernetes cluster has released its second part, focusing on installing the container runtime and core Kubernetes components. The guide covers configuring required Linux kernel modules — overlay and br_netfilter — and setting sysctl parameters to ensure correct network routing between pods and services. Containerd is installed via Docker's official repository using only the containerd.io package, with the systemd cgroup driver enabled to align with kubelet's default configuration. Kubernetes components kubelet and kubeadm are then installed from the official Kubernetes v1.31 repository, with package versions pinned to prevent automatic updates that could break cluster compatibility. All steps are intended to be run on every node in the cluster, both master and workers, unless stated otherwise.

15 AI Crawlers Explained: What Blocking Each One Actually Costs You · ShortSingh