cua-s1-forms
cua-s1-forms is a small tinyx option scorer for GUI form filling, with
706,048 parameters. It belongs to the
Cua-S1 research family.
The Cua-S1 documentation describes it as a tinyx form checkpoint related to
the cua-s1-form-v0 research profile. It is not a formally mapped release of
that profile. For the pinned revision and loader, see
Pinned artifacts.
Load only the safetensors pair. This repository previously shipped a pickle checkpoint,
cua-s1-forms.pt. It has been removed (see trycua/cua#3977). Unpickling a checkpoint can run arbitrary code, so do not load any.ptcopy of this model withtorch.loadorpickle, including copies downloaded from older revisions.cua_s1refuses pickle checkpoints by design.
This model does not generate text. Given a UI element and a list of typed
options, it returns one probability per option in a single forward pass. The
options are one per document entity, plus check, click, and skip. This
is the same input and output contract as TypeSafe's
Jev. The
model scores every actionable element on a form independently, in one
parallel batch. Downstream code, not the model, decides the execution order:
fills, then checkboxes, then the single submit click.
Files
| File | Contents |
|---|---|
cua-s1-forms.safetensors |
Tensors only |
cua-s1-forms.json |
Architecture config, a SHA-256 signature over the tensors, and free-form metadata, including best validation metrics |
This is the only supported format. The pair encodes the same weights as the removed pickle; the converted pair produced bit-for-bit identical model output.
Usage
Install the cua-s1 package from a checkout of
trycua/cua
(uv sync --project libs/cua-s1/python), then run:
from pathlib import Path
from huggingface_hub import snapshot_download
from cua_s1.model import load_checkpoint, select_device
root = Path(snapshot_download(
"cua-ai/cua-s1-forms",
allow_patterns=["cua-s1-forms.safetensors", "cua-s1-forms.json"],
))
# Validates the format, version, and SHA-256 tensor signature before returning.
model, collator, config = load_checkpoint(root / "cua-s1-forms.safetensors", select_device("auto"))
With the CLI, exclude pickle files explicitly when you pin an older revision:
hf download cua-ai/cua-s1-forms --exclude "*.pt" --local-dir ./cua-s1-forms
For the full snapshot, score, order, and execute loop against a live Cua
Driver session, see
cua_s1/planner.py.
The jev-use closed-candidate chooser supports only the Cua-S1-4B adapters and
does not load this checkpoint.
Architecture
- A byte-level embedding and a 2-layer Transformer encoder (width 128, 4 heads) encode the context and, separately, each option's text.
- An option-attention head uses each option as a query against the context tokens to produce an attended context vector. A shared dot product turns each option and attended-context pair into one logit, followed by softmax over the live option count.
- 706,048 trainable parameters.
Input and output
Each element has one context, truncated to 224 bytes:
TASK fill the form from the document, then submit
FORM Northwind Clinic - New Patient Registration
ELEMENT Edit "Phone number" value=""
Options are one per document entity, plus the three fixed actions, each
truncated to 96 bytes: fill Tel: (503) 555-0142, fill DOB: 03/14/1987,
..., check, click, skip.
The model returns one probability per option. The executor picks the argmax.
For a fill, it looks up the entity by index. It then orders the resulting
actions before sending them to Cua Driver (set_value or click).
Training
- 10,000 synthetic episodes from
cua_s1/synth.py. Each episode has a random form of 2 to 16 fields, drawn from a 55-concept catalogue with synonyms for form labels and document labels. Each form is paired with a random fictional person and a random document containing distractor entities and forced look-alike confuser pairs, such asemailandstreet,phoneandemergency contact phone, orstateanduniversity. Window-title suffixes are randomized, and 20% of titles are dropped. - The generator uses reserved phone numbers,
.invaliddomains, invalid test identifiers, and fictional brands. No real user submissions are included. - Splits are disjoint by exact form field signature, so a test form's field set never appears in training.
- AdamW, a cosine schedule with warmup, 6 epochs, batch size 128, and cross-entropy over the live option count.
Results
The checkpoint's author reported these results at publication. The
supporting results document is not part of the trycua/cua repository, and
libs/cua-s1 does not reproduce these results. The best validation metrics
are recorded in cua-s1-forms.json.
| Split | Top-1 | Notes |
|---|---|---|
| Synthetic test (form-disjoint, about 15k decisions) | 99.95% | Hard confuser pairs forced in |
| Real demo eval (3 real forms and 3 real PDFs, 196 decisions, no synthetic data) | 100% | |
| Shuffled-context control | 37% | Indicates that the model reads the element rather than option statistics |
In a head-to-head on the same task, the author also reported 99.7% for this
model and 83.6% for the hosted Jev API (jev-latest, no fine-tuning). Hosted
Jev scored 96% on decisions that require judgment (fill, check, or click). It
scored 74% on recognizing an already-filled field as a no-op, a convention
this model was trained on and hosted Jev was not.
Intended use
- Research on narrowly specified form-oriented computer-use tasks.
- Evaluating specialist-model behavior in isolated, controlled environments.
- Studying task-specific failure modes, verification, and human oversight.
Out of scope
- General-purpose or open-ended computer operation.
- Unsupervised operation on production accounts or sensitive data.
- Actions with financial, legal, medical, employment, safety, or other high-impact consequences.
- Bypassing access controls, consent, rate limits, or service policies.
- Treating model output or apparent task completion as proof that an action was correct or successful.
Limitations
- It chooses only among entities that a document extractor already found as
Label: valuepairs. It cannot invent a value. - It was trained entirely on synthetic forms and evaluated on a small, 196-decision real set. It has not been validated on arbitrary real-world forms outside the demo set.
- It uses a byte-level encoder and an English-centric label vocabulary.
- It is not calibrated with TypeSafe's RLCD method. This is an independent research checkpoint, not a reproduction of Jev.
- No weights-backed CI or canonical Cua Driver desktop E2E covers this checkpoint.
License
MIT.