How Far Should an AI Agent Be Allowed to Go?
An evening for people building, shipping, and scaling with AI. In this hands-on session with Adeuk, we gave an AI agent real tools, increasing autonomy, and a series of increasingly difficult assignments. What happens when it tries to access restricted data? Exceeds its authority? Takes an irreversible action? Finds a creative path around the rules?
- 6:00 PM Networking, food, and some dangerous ideas · meet other builders and propose challenges for the agent
- 6:30 PM The Agent Gauntlet with Adeuk · red-team the agent and the policies, review what executed, what was stopped, and why
- 8:00 PM Surviving production · open discussion
What happened
The Agent Gauntlet: the room pushed an AI agent toward the edge of what it should be allowed to do, then built the controls that keep it from going over. Builders connected an agent to real tools, raised its permissions step by step, and attempted the actions you would be nervous to try in production: restricted data, exceeded authority, irreversible operations, creative paths around the rules. Every allow, deny, log, and escalate decision left an evidence trail the room inspected together. Food and drinks by Adeuk. Photos below.
- Participants connected an agent to real tools and workflows, increased its permissions and autonomy, and attempted actions that were risky, ambiguous, or outside its authority.
- The room wrote policies that allow, deny, log, or escalate agent actions for review, tried to outsmart their own guardrails, and inspected the evidence trail explaining every decision.
Scroll for more, tap a photo to enlarge.
Missed it? Get the next one.
Subscribers get the recap, slides, and the next Boston event before the public Luma drop.
The AI Runtime newsletter
Production lessons from AI practitioners. New issues every week.
Read by AI practitioners at IBM, Amazon, Meta, Google, Nvidia, OpenAI, MIT, Harvard, and more for production learnings.
Free. No spam. Unsubscribe in one click. See past issues →