replica source-backed
JevForge: decision-data synthesis, training and web-action scoring
An end-to-end research toolkit builds decision records, trains and calibrates candidate scorers, and serves a Jev-compatible API. Its released Qwen3.5-0.8B model ranks supplied web actions, while the public demo presents recorded Mind2Web decisions.
Notes
The model card reports Choice top-1 of 0.5787 on 2,400 test records and 0.6373 on 1,158 OOD records using website-disjoint splits and a saved temperature of 0.9717.
source-backed — Public repository, docs or live artifact. About this label