The premise: most prompt work is guesswork because nothing measures it. Forge closes the loop — pick a reference output from the expensive model, then iterate the prompt against the cheap one until parity crosses a threshold you set. The interesting part turned out to be the scoring, not the editing.