smolaya-adfraud

A 323M-parameter decision model that labels an HTTP request fingerprint as human or bot.

Fine-tuned from anon767tom/smolaya — a 20-layer depth-pruned laya — on 22,237 request fingerprints. On the held-out evaluation below it scores 0.778 AUC, against 0.724 for the base model on the identical task, labels and splits. It runs on CPU and answers in a single forward pass. A smaller model trained on a small subset of the production logs of https://nobotspls.com.

Usage

With transformers:

from transformers import pipeline

clf = pipeline("zero-shot-classification",
               model="anon767tom/smolaya-adfraud", trust_remote_code=True)

Q = {"adfraud": {
    "type": "choice",
    "instructions": "Who sent this HTTP request?",
    "criteria": {
        "human": "a real person using a normal consumer web browser",
        "bot": "a bot, crawler, scraper, headless browser, or automated traffic "
               "routed through a proxy, VPN or datacenter"}}}

state = """HTTP request fingerprint.
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/15.6 Safari/605.1.15
Sec-CH-UA: "Not A;Brand";v="99", "Chromium";v="118"
Accept-Language: en-GB,en;q=0.8
Accept-Encoding: gzip, deflate, br
Header set: host, user-agent, accept-language, accept-encoding, accept
Client IP: 198.51.100.77"""

print(clf(state, questions=Q))
# {'adfraud': {'type': 'choice', 'choice': 'bot',
#              'probabilities': {'human': 0.0003, 'bot': 0.9997}, ...}}

That fingerprint pairs a Safari 15.6 User-Agent with a Chromium 118 Sec-CH-UA — the kind of cross-field inconsistency the model picks up.

With the laya package:

import laya

ag = laya.load("anon767tom/smolaya-adfraud", device="cpu")
out = ag.predict(state, {"a": Q["adfraud"]})["answers"]["a"]
print(out["choice"], out["probabilities"])

A fingerprint the model scores the other way — a genuine Edge 138 browser session:

HTTP request fingerprint.
User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/138.0.0.0 Safari/537.36 Edg/138.0.0.0
Sec-CH-UA: "Not)A;Brand";v="8", "Chromium";v="138", "Microsoft Edge";v="138"
Accept-Language: en-US,en;q=0.9
Accept-Encoding: gzip, deflate, br, zstd
Header set: content-length,cache-control,sec-ch-ua-platform,user-agent,sec-ch-ua,content-type,sec-ch-ua-mobile,accept,origin,sec-fetch-site,sec-fetch-mode,sec-fetch-dest,referer,accept-encoding,accept-language,priority
Client IP: 203.0.113.42

-> human {'human': 0.9951, 'bot': 0.0049}

Both examples use RFC 5737 documentation addresses in place of the original client IPs. The question above is the one the model was trained on; changing its wording or the option descriptions changes the scores.

Input format

Six fields, in this order:

HTTP request fingerprint.
User-Agent: <value>
Sec-CH-UA: <value>
Accept-Language: <value>
Accept-Encoding: <value>
Header set: <comma-separated header names>
Client IP: <value>

Each field is in one of three states, and the difference matters:

state how to write it meaning
observed Accept-Language: en-GB the client sent this value
absent Accept-Language: (absent) the client sent no such header — a real signal
unobservable omit the line entirely your vantage point cannot see this field

Do not substitute a placeholder for an unobservable field; omit the line. The model was trained with this convention and treats an omitted line and (absent) differently.

Header set is the set of header names. Most training rows come from a deployment behind a Go reverse proxy that re-emits headers with host and user-agent first and the rest sorted, and that injects ja3, x-forwarded-for, x-http-version, x-ja3-hash and x-sec-fetch-*. The model reads the fingerprint as a whole rather than field by field, so feed it real captures in your own edge's convention rather than hand-assembled inputs.

Training data

22,237 fingerprints from four sources. Rows from the production source carry 8× the sample weight of any other row; class balance is then computed on the weighted set so the upweighting does not shift the prior.

source rows label fields observed
Undisclosed production dataset (weighted 8×) 1,838 551 human / 1,287 bot all six
FP-Stalker 15,000 human no Sec-CH-UA, no IP
crawler-user-agents + well-known-bots 2,177 bot (declared crawlers) User-Agent only
intoli/user-agents 969 human User-Agent only
Self-captured header endpoint 2,253 2,249 bot / 4 human no Accept-Encoding, no IP

Open-source sources, in full

These are cited as repositories rather than in the datasets: metadata field because none of them is published as a Hugging Face dataset.

  • FP-Stalker — Spirals-Team/FPStalker. 15,000 fingerprints sampled from 1,819 browser instances, collected via the AmIUnique browser extensions between July 2015 and August 2017. Real browsers driven by real people. From Vastel, Laperdrix, Rudametkin and Rouvoy, FP-STALKER: Tracking Browser Fingerprint Evolutions, IEEE S&P 2018. Used as the human class; the HTTP header columns only (userAgentHttp, languageHttp, encodingHttp, orderHttp), with the JavaScript and Flash fingerprint columns discarded.
  • crawler-user-agents — monperrus/crawler-user-agents, MIT. Declared-crawler User-Agent strings.
  • well-known-bots — arcjet/well-known-bots, Apache-2.0. Declared-bot User-Agent strings.
  • intoli/user-agents — intoli/user-agents, MIT. 969 unique User-Agent strings from real visitor profiles, used as the human side at matching observability.

The undisclosed production dataset is not distributed.

Training recipe

Encoder layers 12-19, final_norm, the type embedding and the decision head were trained — 124.3M of 323.2M parameters; layers 0-11 and the embeddings stayed frozen at the smolaya initialisation. Two-option choice objective with the option order randomised every step to remove position bias, 1 epoch over 22,237 rows, batch 16, AdamW at 2e-5 (encoder) / 5e-5 (head), linear decay with warmup, fp16 autocast, 1,389 steps at 0.47 s/step on one Tesla T4.

Evaluation

5-fold GroupKFold grouped by User-Agent over the 1,838 production rows, so no User-Agent seen in training reappears in a test fold. Every row from every other source trains in every fold, except an external row whose User-Agent occurs in the fold under test, which is held out of that fold only. Scores are out-of-fold; the interval is a User-Agent-cluster bootstrap.

metric value
ROC AUC 0.778 [0.643, 0.916]
Average precision 0.896
Brier score 0.236
Base smolaya, same task and splits 0.724
Per-fold AUC 0.579 · 0.905 · 0.929 · 0.828 · 0.606

At a 0.5 threshold on the same out-of-fold predictions: flags 73% of requests at precision 0.802, recall 0.840. Thresholds of 0.3 and 0.7 move precision by under a point — the score distribution is concentrated near 0 and 1, so fit isotonic regression on your own rows before reading the outputs as probabilities.

Two results that bear on how the numbers should be read:

  • The production rows are labelled with that system's own calibrated decisions, not human-verified ground truth. Every figure above measures agreement with an existing production classifier and by construction cannot exceed it.
  • On production rows where three or more rules fired — overwhelmingly IP-reputation blocks on requests whose headers look like an ordinary browser — AUC is 0.09, i.e. strongly inverted: the model calls them human. It sees no IP reputation, velocity or ASN data, and header text does not substitute for a list lookup. Pair it with those signals rather than replacing them.

Latency

fp32, CPU, 4 threads, batch 1, 165 input tokens: mean 1.096 s, p50 1.048 s, p95 1.439 s per request. Dynamic int8 quantisation (see anon767tom/smolaya-int8, which skips mlp.Wo to avoid activation outliers) gives roughly 2× on CPU. Suited to offline scoring, sampled auditing or a second-stage check rather than inline blocking at the edge.

License

Apache-2.0, inherited from laya. The public sources retain their own licenses as listed above.

Downloads last month
22
Safetensors
Model size
0.3B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for anon767tom/smolaya-adfraud

Finetuned
(2)
this model