replica source-backed
JEVfire: parallel finite-choice decisions with vLLM
An independent CUDA runtime scores verified single-token option labels with existing language-model weights. It batches independent fields through vLLM, reuses shared prefixes when supported and assembles typed output in code; browser experiments also use local WebGPU models.
Notes
On the author's synthetic 28-field task with Qwen3.8-27B-FP8 and an RTX PRO 6000 Blackwell, fresh-prefix median latency was 496.9 ms versus 5,113.1 ms for constrained JSON, with five trials per cell.
source-backed — Public repository, docs or live artifact. About this label
Destinations
Availability & demo checks
HTTP 200 · HTTP response only. Checked destination ↗ · Last reachable