proof
we planted the answers, then asked the engine.
Every dataset below is synthetic with a property planted by construction, so the engine's verdict can be marked right or wrong rather than plausible. These are the unedited results, contracts and all, from the same build that ships on AWS Marketplace.
the full chain
what we planted
+6/week trend with a +25 jump at week 18, both planted
what we asked
“how are signups doing”
what came back
Yes: signups shifted. The instruments agree the mean moved (tripwire fired at row 29).
smooth on signups: the instruments returned no actionable read on this data.
signups is rising at +5.0000 per step; paper only until the move clears the error bar.
detect_change: ACT smooth: None trend: PAPER_ONLY
the receipt
11/11 solver contracts ok · certificates: input_manifest · solve 9,014 ms
download this dataset and run it against your own instance.
trend
what we planted
grows ~6%/month by construction; wobble sums to zero
what we asked
“is revenue actually growing month over month”
what came back
revenue is rising at +3426.5000 per step; the evidence clears the bar to act.
trend: ACT
the receipt
4/4 solver contracts ok · certificates: input_manifest · solve 3,247 ms
download this dataset and run it against your own instance.
change detection
what we planted
mean steps 50 -> 58 at row 30 by construction
what we asked
“did this sensor shift and when”
what came back
Yes: temp_c shifted. The instruments agree the mean moved (tripwire fired at row 47).
detect_change: ACT
the receipt
5/5 solver contracts ok · certificates: input_manifest · solve 4,415 ms
download this dataset and run it against your own instance.
smoothing
what we planted
flat 120 by construction; the saw sums to zero
what we asked
“what is the underlying demand under all this noise”
what came back
smooth on demand: the instruments returned no actionable read on this data.
smooth: None
the receipt
2/2 solver contracts ok · certificates: input_manifest · solve 1,607 ms
download this dataset and run it against your own instance.
ranking
what we planted
eyelet 44.0, anchor 41.0, crank 38.5 by construction
what we asked
“which three products carry the best margin”
what came back
rank on margin: the instruments returned no actionable read on this data.
rank: None
the receipt
3/3 solver contracts ok · certificates: input_manifest · solve 2,456 ms
download this dataset and run it against your own instance.
distribution
what we planted
40 fast calls near 43ms, 3 planted stragglers near 195ms
what we asked
“what does the response time distribution look like”
what came back
distribution on ms: the instruments returned no actionable read on this data.
distribution: None
the receipt
2/3 solver contracts ok · certificates: input_manifest · solve 2,417 ms
download this dataset and run it against your own instance.
routing
what we planted
routing should pick change detection; mean 12 -> 31 at row 20
what we asked
“did the average error rate shift after the rollout”
what came back
Yes: errors shifted. The instruments agree the mean moved (tripwire fired at row 35).
detect_change: ACT
the receipt
5/5 solver contracts ok · certificates: input_manifest · solve 4,062 ms
download this dataset and run it against your own instance.
Run the same check yourself: subscribe, launch, and hand the console a problem you
already know the answer to. If a verdict here ever reads PAPER_ONLY,
that is the gate declining to certify a thin result, which is the product working.