How a run works

  1. 01

    You write the rule.

    Plain language, exactly as it would ship, plus the one event that proves the business got what it paid for.

  2. 02

    Four challengers try to break it.

    One walks the journey you intended. One looks for the fewest steps that still collect the reward. One tries to collect it again. One checks whether your written rule gives anyone the authority to stop it. They run one at a time, against your own text.

  3. 03

    You get a counterexample, or you get told there isn't one.

    A run that receives the reward without ever proving value stops the bench and shows the shortest path to it. If no challenger finds one, the verdict says so. The product never manufactures a finding.

The challengers

Four classes run today, one at a time, each against the exact text you submitted. The code beside a name is how you tell its lane apart inside a run.

A1 EXPECTED
Completes the journey you intended and collects the reward the way you meant it to be collected. It is the control for the other three: if this one cannot get paid, the rule is broken in the opposite direction.
A2 SHORTEST PATH
Receives the reward with the fewest possible actions. This is the class that finds loopholes, and when it succeeds without reaching proof of value, its route becomes the shortest counterexample.
A3 REPEAT
Repeats the shortest successful path, or coordinates it. A rule can be sound once and unsound in volume, and the difference is what turns a curiosity into an exposure.
A4 CONTROL
Tries to stop the shortcut using only the powers your written rule actually grants: a cap, an approval, a verification, a reversal. A rule that grants none permits everything it does not mention.

The honesty rules are the product.

An incentive review that overstates itself is worse than none, because it gets believed. These are load-bearing, not marketing.

Adversarial tests, not predictions.
MISBEHAVE generates plausible ways a rule can be satisfied without value being created. It does not forecast what real people will do, and it never claims to.
Permitted exposure, not predicted loss.
The exposure figure is arithmetic on the rule's own number, scaled over 1, 10 or 100 events. It says what the written text mechanically permits. It says nothing about likelihood.
No invented numbers.
No probabilities, no confidence percentages, no exploitability score, no gauges. Only counts that trace back to a generated test.
Never “safe”.
The best verdict this product will give you is “passing under the current test suite”, because four tests cannot establish anything stronger.
Every quotation is verified.
A cited phrase is checked as an exact substring of the rule you submitted before it is ever displayed.
Repairs show their cost.
Each guardrail carries its tradeoff. A patch is never presented as free.

What this product cannot tell you

These are refusals, not gaps waiting to be filled. A review that overstates itself is worse than none, because it gets believed.

Whether anyone will actually do this.
A challenger is a strategy against your text, not a forecast of a person. The product generates plausible ways a rule can be satisfied without value being created, and it never claims to know what real people will choose.
How much you will lose.
The exposure figure is arithmetic on your own number, repeated 1, 10 or 100 times. It is what the written text mechanically permits, and it is not a loss estimate, a probability or a forecast.
That your rule is safe.
The best verdict here is passing under the current test suite. Four classes cannot establish more than that, so the product does not say more than that.
Whether your rule is lawful.
It reads incentive logic, not regulation. Nothing here is legal, tax or accounting advice, and a clean verdict says nothing about compliance in any jurisdiction.