A prompt-set benchmark groups realistic customer questions and checks them consistently across runs or search environments. It avoids treating a couple of examples as the whole picture.
The demo runs three questions over three rounds and draws their movement. It exposes variability that an average can hide.
Keep the prompt list, region, language, date, and scoring rule together. Answers may vary, so avoid declaring a winner from one run.
When to use
Use it to track GEO work over time while retaining answer variability.