A written IAM access policy (5 roles, 6 departments, 10 named systems), a rules engine
that encodes it, and 90 test scenarios checking whether an LLM applies the policy the
way a human reviewer would. Try the simulator below, or scroll down for the eval results.
Full write-up on GitHub.
Try it yourself
Pick a scenario and see what the policy says, instantly. This runs entirely in your
browser against the same rules engine used to generate the 90 test cases below, no
API key or model call involved.
Eval results
Results pending: this page ships with all 90 generated scenarios and their expected
answers, but the model hasn't been run against them yet on this page. Once the eval runs,
this section fills in with pass/fail and the model's actual output for every case.