Classify HTTP User-Agent strings to determine if traffic is human, bot, AI/LLM crawler, or AI agent. Detects headless browsers, scores aggressiveness, and verifies bot identity via IP.
Every request passes through 11 detection layers in under 10ms. The system combines deterministic rules with machine learning — rules catch known threats instantly, ML catches novel ones that rules miss.
{ isBot, category, intent, confidence, signals[], mlScore }
Think of it like airport security with multiple checkpoints. First, we check your passport (User-Agent string) against a watchlist of 1650 known bots. Then we verify your boarding pass (IP address) matches who you claim to be. Finally, an AI system looks at your overall behavior pattern — even if you have a perfect fake passport, the AI might notice something off.
| Layer | Method | Latency | What it catches |
|---|---|---|---|
| Pattern Matching | Regex against classifier-db.json | <1ms | Known bots (Googlebot, GPTBot, scrapers) |
| Headless Detection | UA substrings + Client Hints analysis | <1ms | HeadlessChrome, Puppeteer, PhantomJS |
| AI Agent Detection | Pattern list (10 known agents) | <1ms | browser-use, Operator, AgentQL, Devin |
| IP → ASN | MaxMind MMDB binary search | <1ms | Network operator identification |
| Bot IP Verification | CIDR range matching | <1ms | Spoofed vs. genuine bots |
| Datacenter / Hosting | CIDR (AWS/GCP/Azure) + ASN set (hosting/colo) | <1ms | Cloud, hosting & colo traffic (Alibaba, Tencent, OVH, Hetzner) |
| IP Reputation | Set lookup (17k+ IPs from 30+ blacklists) | <1ms | Known malicious IPs, scanners, brute-force |
| Reverse DNS | Cached PTR lookup + domain verify | 0-1ms | Hostname-based bot verification |
| ML Model | LightGBM via ONNX Runtime (634 features) | <5ms | Novel bots, structural anomalies |
| Behavioral Scoring | Client signals: CDP, canvas hash, RTT, plugins, interaction | <1ms | Headless browsers with real Chrome UA |
| Confidence | Multi-signal fusion | <1ms | Final verdict with explainability |
The machine learning model trains daily on real production traffic. It learns patterns that humans can't write rules for — like the fact that a 47-character User-Agent with 3 slashes, no parentheses, and the trigram "bot" at position 12 is a 94% bot indicator. The model currently achieves 99.8% AUC-ROC and improves automatically every day.
The ML model acts as an auditor on top of the deterministic rules: it can reclaim borderline no-evidence bots back to human when its score is very low, escalate a human verdict to bot when its score is very high (>0.95) and a hard fact corroborates it (a malicious-IP reputation hit, or a literal automation string in the User-Agent itself: HeadlessChrome, PhantomJS, Puppeteer), and reinforce confidence when it agrees with the rules. Hard facts — a verified bot IP, a malicious-IP or honeypot hit, or a known bot name — always win and are never overridden.
Two categories of evidence are deliberately excluded from ever corroborating that escalation, even combined:
client_hints_incomplete, missing_client_hints_for_modern_chrome). A corporate proxy or security gateway stripping that header can trigger these on a genuine human; only a literal automation string in the UA counts as hard evidence.Separately, any established positive human evidence (Siemens corporate egress, or consent corroborated by genuine interaction) is an absolute veto on the ML escalation: once a request is proven human, the model cannot override it at any score.
This is how the API classifies you (the current visitor):
curl -X POST https://ua-api.lab.c2comms.cloud/classify \
-H "Content-Type: application/json" \
-d '{
"userAgent": "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)",
"ip": "216.73.216.136"
}'
Classifying you…
Enter any User-Agent string to classify it in real-time:
Click "Classify" to see results
| Source | Purpose | License |
|---|---|---|
| arcjet/well-known-bots | Primary bot DB (600+ bots) | Apache 2.0 |
| monperrus/crawler-user-agents | Additional regex patterns | MIT |
| ai-robots-txt/ai.robots.txt | AI/LLM bot identification | MIT |
| Google / Bing / OpenAI | Bot IP verification ranges | Public |
| ip-location-db ASN MMDB | IP → ASN lookup | CC0 |
| stamparm/ipsum | IP reputation (17k+ malicious IPs from 30+ blacklists) | Unlicense |
| brianhama/bad-asn-list | Datacenter/hosting ASN set (drives ua_hosting, + Tencent extras) | MIT |
This page runs the same fingerprinting script that's deployed on siemens.com. It collects 25+ signals from your browser environment and interaction patterns. A real human typically scores 0-10. A headless bot scores 60+.
Collecting signals (2.5s)...
Signals: webdriver, plugins, mimeTypes, GPU renderer, canvas hash, CDP detection (cdc_ vars + Runtime serialization probe), network RTT, document focus, notification permission + Permissions API mismatch, mouse/scroll/touch/keyboard interaction, time-to-first-interaction.