You don't need a complete policy to start
The instinct is to sit down and enumerate every action an agent might take, decide allow/hold/deny for each one and only then let it near anything real. That's backwards for the same reason it's backwards for human access control: you don't write a complete permission matrix before hiring your first employee. You start with sensible defaults for the categories that matter and you tighten specific cases as you learn where the actual risk is.
A working starting policy needs three things, not a spreadsheet of every possible action.
1. Sort actions by consequence, not by type
"Read" versus "write" is a reasonable first cut, but it's not the axis that actually matters. A write to a staging database and a write to a production payments table are both "writes" and have almost nothing else in common. Sort by what happens if the action is wrong instead:
Reversible, low blast radius - reading records, querying internal tools, drafting content nobody sees yet. Allow by default.
Reversible, but visible or costly to undo - sending a customer-facing message, modifying a non-production environment, issuing a small refund. Allow within limits; log everything.
Hard to reverse, or expensive to get wrong - production deploys, large financial transactions, data exports, anything touching customer PII at scale. Hold for a human by default, every time, no exceptions carved out later under deadline pressure.
2. Attach every action to a target, not just a verb
"Can this agent deploy?" is not answerable on its own. "Can this agent deploy to staging?" and "can this agent deploy to production?" are different questions with different answers. The same verb, scoped to a different target, belongs in a different consequence tier. Write policy against (action, target) pairs, not actions alone - it's more work up front and it's the difference between a policy that actually holds and one that technically exists.
3. Decide who holds the actions you don't auto-allow
"Requires approval" isn't a complete answer if nobody knows whose approval. Route by context that already exists in your organization - the on-call engineer for deploys, the account owner for large refunds, a data steward for exports - rather than a single generic approver queue that becomes a bottleneck nobody reads carefully. A held action that takes two days to clear because it landed on the wrong desk trains people to route around the policy entirely.
Start narrow, expand deliberately
The practical path isn't "write the complete policy, then deploy the agent." It's: pick one agent and one real workflow, write policy for the handful of actions it actually takes, deploy it with the third tier - hard to reverse, expensive to get wrong - held for approval by default and expand coverage as the agent's scope genuinely grows. A narrow, enforced policy beats a comprehensive one that exists only as a document.

