User Agent Classifier API

Classify HTTP User-Agent strings to determine if traffic is human, bot, AI/LLM crawler, or AI agent. Detects headless browsers, scores aggressiveness, and verifies bot identity via IP.

Known Bots

1650
Patterns in classifier database

AI/LLM Crawlers

203
GPTBot, ClaudeBot, CCBot, etc.

Categories

17
Including ai_agent for browser automation

Detection Layers

11
Pattern + ML + IP + behavioral + intent

How it works

Every request passes through 11 detection layers in under 10ms. The system combines deterministic rules with machine learning — rules catch known threats instantly, ML catches novel ones that rules miss.

The Pipeline (simplified)

① User-Agent string → Pattern match against 1650 known bots
② Headless check → HeadlessChrome? PhantomJS? Missing Client Hints?
③ AI Agent check → browser-use? Operator? AgentQL?
④ IP address → ASN lookup → Who owns this network?
⑤ Bot verification → Does the IP match Google/Bing/OpenAI published ranges?
⑥ Datacenter / hosting → AWS/GCP/Azure CIDRs + hosting/colo ASNs (Alibaba, Tencent, OVH, Hetzner, …)?
⑦ IP reputation → Is this IP on 3+ threat intelligence blacklists?
⑧ Reverse DNS → Does the hostname confirm the bot's identity?
⑨ ML model → LightGBM scores 0.0 (human) to 1.0 (bot)
⑩ Behavioral signals → webdriver? CDP? Canvas hash? No plugins? No interaction?
⑪ Confidence → Combine all signals → high / medium / low
→ Result: { isBot, category, intent, confidence, signals[], mlScore }

For non-technical users

Think of it like airport security with multiple checkpoints. First, we check your passport (User-Agent string) against a watchlist of 1650 known bots. Then we verify your boarding pass (IP address) matches who you claim to be. Finally, an AI system looks at your overall behavior pattern — even if you have a perfect fake passport, the AI might notice something off.

For engineers

LayerMethodLatencyWhat it catches
Pattern MatchingRegex against classifier-db.json<1msKnown bots (Googlebot, GPTBot, scrapers)
Headless DetectionUA substrings + Client Hints analysis<1msHeadlessChrome, Puppeteer, PhantomJS
AI Agent DetectionPattern list (10 known agents)<1msbrowser-use, Operator, AgentQL, Devin
IP → ASNMaxMind MMDB binary search<1msNetwork operator identification
Bot IP VerificationCIDR range matching<1msSpoofed vs. genuine bots
Datacenter / HostingCIDR (AWS/GCP/Azure) + ASN set (hosting/colo)<1msCloud, hosting & colo traffic (Alibaba, Tencent, OVH, Hetzner)
IP ReputationSet lookup (17k+ IPs from 30+ blacklists)<1msKnown malicious IPs, scanners, brute-force
Reverse DNSCached PTR lookup + domain verify0-1msHostname-based bot verification
ML ModelLightGBM via ONNX Runtime (634 features)<5msNovel bots, structural anomalies
Behavioral ScoringClient signals: CDP, canvas hash, RTT, plugins, interaction<1msHeadless browsers with real Chrome UA
ConfidenceMulti-signal fusion<1msFinal verdict with explainability

The ML advantage

The machine learning model trains daily on real production traffic. It learns patterns that humans can't write rules for — like the fact that a 47-character User-Agent with 3 slashes, no parentheses, and the trigram "bot" at position 12 is a 94% bot indicator. The model currently achieves 99.8% AUC-ROC and improves automatically every day.

The ML model acts as an auditor on top of the deterministic rules: it can reclaim borderline no-evidence bots back to human when its score is very low, escalate a human verdict to bot when its score is very high (>0.95) and a hard fact corroborates it (a malicious-IP reputation hit, or a literal automation string in the User-Agent itself: HeadlessChrome, PhantomJS, Puppeteer), and reinforce confidence when it agrees with the rules. Hard facts — a verified bot IP, a malicious-IP or honeypot hit, or a known bot name — always win and are never overridden.

Two categories of evidence are deliberately excluded from ever corroborating that escalation, even combined:

Separately, any established positive human evidence (Siemens corporate egress, or consent corroborated by genuine interaction) is an absolute veto on the ML escalation: once a request is proven human, the model cannot override it at any score.

Your classification

This is how the API classifies you (the current visitor):

curl -X POST https://ua-api.lab.c2comms.cloud/classify \
  -H "Content-Type: application/json" \
  -d '{
    "userAgent": "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)",
    "ip": "216.73.216.136"
  }'
Classifying you…

Try it live

Enter any User-Agent string to classify it in real-time:

Click "Classify" to see results

Categories

advertising ai archiver content_fetcher feed_fetcher http_library internal_service link_checker monitoring page_preview scraping_framework search_engine security seo social_media tool webhooks

Data sources

SourcePurposeLicense
arcjet/well-known-botsPrimary bot DB (600+ bots)Apache 2.0
monperrus/crawler-user-agentsAdditional regex patternsMIT
ai-robots-txt/ai.robots.txtAI/LLM bot identificationMIT
Google / Bing / OpenAIBot IP verification rangesPublic
ip-location-db ASN MMDBIP → ASN lookupCC0
stamparm/ipsumIP reputation (17k+ malicious IPs from 30+ blacklists)Unlicense
brianhama/bad-asn-listDatacenter/hosting ASN set (drives ua_hosting, + Tencent extras)MIT

Bot Signals (live from your browser)

This page runs the same fingerprinting script that's deployed on siemens.com. It collects 25+ signals from your browser environment and interaction patterns. A real human typically scores 0-10. A headless bot scores 60+.

Collecting signals (2.5s)...

Signals: webdriver, plugins, mimeTypes, GPU renderer, canvas hash, CDP detection (cdc_ vars + Runtime serialization probe), network RTT, document focus, notification permission + Permissions API mismatch, mouse/scroll/touch/keyboard interaction, time-to-first-interaction.