Instructions to use anon767tom/smolaya-adfraud with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use anon767tom/smolaya-adfraud with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("zero-shot-classification", model="anon767tom/smolaya-adfraud", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("anon767tom/smolaya-adfraud", trust_remote_code=True, device_map="auto") - Laya
How to use anon767tom/smolaya-adfraud with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
smolaya-adfraud
A 323M-parameter decision model that labels an HTTP request fingerprint as human or bot.
Fine-tuned from anon767tom/smolaya — a 20-layer
depth-pruned laya — on 22,237 request fingerprints.
On the held-out evaluation below it scores 0.778 AUC, against 0.724 for the base model on the
identical task, labels and splits. It runs on CPU and answers in a single forward pass.
A smaller model trained on a small subset of the production logs of https://nobotspls.com.
Usage
With transformers:
from transformers import pipeline
clf = pipeline("zero-shot-classification",
model="anon767tom/smolaya-adfraud", trust_remote_code=True)
Q = {"adfraud": {
"type": "choice",
"instructions": "Who sent this HTTP request?",
"criteria": {
"human": "a real person using a normal consumer web browser",
"bot": "a bot, crawler, scraper, headless browser, or automated traffic "
"routed through a proxy, VPN or datacenter"}}}
state = """HTTP request fingerprint.
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/15.6 Safari/605.1.15
Sec-CH-UA: "Not A;Brand";v="99", "Chromium";v="118"
Accept-Language: en-GB,en;q=0.8
Accept-Encoding: gzip, deflate, br
Header set: host, user-agent, accept-language, accept-encoding, accept
Client IP: 198.51.100.77"""
print(clf(state, questions=Q))
# {'adfraud': {'type': 'choice', 'choice': 'bot',
# 'probabilities': {'human': 0.0003, 'bot': 0.9997}, ...}}
That fingerprint pairs a Safari 15.6 User-Agent with a Chromium 118 Sec-CH-UA — the kind of
cross-field inconsistency the model picks up.
With the laya package:
import laya
ag = laya.load("anon767tom/smolaya-adfraud", device="cpu")
out = ag.predict(state, {"a": Q["adfraud"]})["answers"]["a"]
print(out["choice"], out["probabilities"])
A fingerprint the model scores the other way — a genuine Edge 138 browser session:
HTTP request fingerprint.
User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/138.0.0.0 Safari/537.36 Edg/138.0.0.0
Sec-CH-UA: "Not)A;Brand";v="8", "Chromium";v="138", "Microsoft Edge";v="138"
Accept-Language: en-US,en;q=0.9
Accept-Encoding: gzip, deflate, br, zstd
Header set: content-length,cache-control,sec-ch-ua-platform,user-agent,sec-ch-ua,content-type,sec-ch-ua-mobile,accept,origin,sec-fetch-site,sec-fetch-mode,sec-fetch-dest,referer,accept-encoding,accept-language,priority
Client IP: 203.0.113.42
-> human {'human': 0.9951, 'bot': 0.0049}
Both examples use RFC 5737 documentation addresses in place of the original client IPs. The question above is the one the model was trained on; changing its wording or the option descriptions changes the scores.
Input format
Six fields, in this order:
HTTP request fingerprint.
User-Agent: <value>
Sec-CH-UA: <value>
Accept-Language: <value>
Accept-Encoding: <value>
Header set: <comma-separated header names>
Client IP: <value>
Each field is in one of three states, and the difference matters:
| state | how to write it | meaning |
|---|---|---|
| observed | Accept-Language: en-GB |
the client sent this value |
| absent | Accept-Language: (absent) |
the client sent no such header — a real signal |
| unobservable | omit the line entirely | your vantage point cannot see this field |
Do not substitute a placeholder for an unobservable field; omit the line. The model was trained with
this convention and treats an omitted line and (absent) differently.
Header set is the set of header names. Most training rows come from a deployment behind a Go
reverse proxy that re-emits headers with host and user-agent first and the rest sorted, and that
injects ja3, x-forwarded-for, x-http-version, x-ja3-hash and x-sec-fetch-*. The model reads
the fingerprint as a whole rather than field by field, so feed it real captures in your own edge's
convention rather than hand-assembled inputs.
Training data
22,237 fingerprints from four sources. Rows from the production source carry 8× the sample weight of any other row; class balance is then computed on the weighted set so the upweighting does not shift the prior.
| source | rows | label | fields observed |
|---|---|---|---|
| Undisclosed production dataset (weighted 8×) | 1,838 | 551 human / 1,287 bot | all six |
| FP-Stalker | 15,000 | human | no Sec-CH-UA, no IP |
| crawler-user-agents + well-known-bots | 2,177 | bot (declared crawlers) | User-Agent only |
| intoli/user-agents | 969 | human | User-Agent only |
| Self-captured header endpoint | 2,253 | 2,249 bot / 4 human | no Accept-Encoding, no IP |
Open-source sources, in full
These are cited as repositories rather than in the datasets: metadata field because none of them is
published as a Hugging Face dataset.
- FP-Stalker — Spirals-Team/FPStalker. 15,000
fingerprints sampled from 1,819 browser instances, collected via the AmIUnique browser extensions
between July 2015 and August 2017. Real browsers driven by real people. From Vastel, Laperdrix,
Rudametkin and Rouvoy, FP-STALKER: Tracking Browser Fingerprint Evolutions, IEEE S&P 2018. Used
as the human class; the HTTP header columns only (
userAgentHttp,languageHttp,encodingHttp,orderHttp), with the JavaScript and Flash fingerprint columns discarded. - crawler-user-agents — monperrus/crawler-user-agents, MIT. Declared-crawler User-Agent strings.
- well-known-bots — arcjet/well-known-bots, Apache-2.0. Declared-bot User-Agent strings.
- intoli/user-agents — intoli/user-agents, MIT. 969 unique User-Agent strings from real visitor profiles, used as the human side at matching observability.
The undisclosed production dataset is not distributed.
Training recipe
Encoder layers 12-19, final_norm, the type embedding and the decision head were trained —
124.3M of 323.2M parameters; layers 0-11 and the embeddings stayed frozen at the smolaya
initialisation. Two-option choice objective with the option order randomised every step to remove
position bias, 1 epoch over 22,237 rows, batch 16, AdamW at 2e-5 (encoder) / 5e-5 (head), linear
decay with warmup, fp16 autocast, 1,389 steps at 0.47 s/step on one Tesla T4.
Evaluation
5-fold GroupKFold grouped by User-Agent over the 1,838 production rows, so no User-Agent seen
in training reappears in a test fold. Every row from every other source trains in every fold, except
an external row whose User-Agent occurs in the fold under test, which is held out of that fold only.
Scores are out-of-fold; the interval is a User-Agent-cluster bootstrap.
| metric | value |
|---|---|
| ROC AUC | 0.778 [0.643, 0.916] |
| Average precision | 0.896 |
| Brier score | 0.236 |
| Base smolaya, same task and splits | 0.724 |
| Per-fold AUC | 0.579 · 0.905 · 0.929 · 0.828 · 0.606 |
At a 0.5 threshold on the same out-of-fold predictions: flags 73% of requests at precision 0.802, recall 0.840. Thresholds of 0.3 and 0.7 move precision by under a point — the score distribution is concentrated near 0 and 1, so fit isotonic regression on your own rows before reading the outputs as probabilities.
Two results that bear on how the numbers should be read:
- The production rows are labelled with that system's own calibrated decisions, not human-verified ground truth. Every figure above measures agreement with an existing production classifier and by construction cannot exceed it.
- On production rows where three or more rules fired — overwhelmingly IP-reputation blocks on requests whose headers look like an ordinary browser — AUC is 0.09, i.e. strongly inverted: the model calls them human. It sees no IP reputation, velocity or ASN data, and header text does not substitute for a list lookup. Pair it with those signals rather than replacing them.
Latency
fp32, CPU, 4 threads, batch 1, 165 input tokens: mean 1.096 s, p50 1.048 s, p95 1.439 s per
request. Dynamic int8 quantisation (see anon767tom/smolaya-int8,
which skips mlp.Wo to avoid activation outliers) gives roughly 2× on CPU. Suited to offline
scoring, sampled auditing or a second-stage check rather than inline blocking at the edge.
License
Apache-2.0, inherited from laya. The public sources retain their own licenses as listed above.
- Downloads last month
- 22