How a document analytics tool learned to tell real readers from bots
A document-analytics startup that tracks who reads shared PDFs found its viewer counts were being inflated by automated systems including corporate email security scanners, link-preview bots from Slack and LinkedIn, and uptime monitors. To filter these out, the team built a weighted scoring system that flags sessions as suspicious when bot-like signals — such as missing user agents, zero engagement, or sub-second page views — cross a threshold of 60 points. The process revealed unexpected edge cases, including the Android phone brand CUBOT whose device name contains the substring 'bot', causing real users to be misclassified as crawlers. The team also discovered that a 1-millisecond dwell-time floor was being triggered by server-side processes rather than genuine human interaction, leading them to raise the minimum plausible reading duration to 10 seconds. The project highlighted that adding evidence meant to exonerate real users sometimes backfired, making legitimate sessions appear even more bot-like due to inverted scoring logic.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in