Skip to content
Bentley PerkinsAn independent commonsFuturism Institute

Bentley Perkins · applied AI systems

AI workflows that show what they did.

I help small teams turn repetitive research, proposal, intake and reporting work into controlled systems, with evidence, permissions and human approval built in. Then I write down what the system cannot do, and hand you the instructions for turning it off.

Seattle and remote · two founding-client engagements open, as of August 2026

Refusal and escalationAI assistanceDeterministic checksHumanapproval

What I mean by controlled

Human approval
A person sees the evidence and decides. Nothing consequential passes this without them.
Deterministic checks
Code, not judgement. Schema, arithmetic, source presence, permission.
AI assistance
Drafting, extraction, comparison, summary. Never the final word on anything.
Refusal and escalation
Out of scope, low evidence, or untrusted input. The system stops and says why.

Most AI workflows have the middle ring and nothing else. The work is the other three.

Four ways to work together, and what each costs

Fixed scope and a named price. If a job does not fit one of these, I will say so rather than stretch one to cover it.

Four engagements, what each runs, and what it costs
EngagementRunsPrice
Evidence-First Workflow Sprintten working days$3,000founding rate, first two clients
Agent Reliability Audittwo weeks$4,500founding rate
AI Workbench Clinic90 minutes$300per session
Managed Stewardshipmonthly, cancel any time$500 to $1,000per month

Evidence-First Workflow Sprint

$3,000 founding rate, first two clients

One repetitive process becomes a controlled system. The main offer, and the one to read first.

  • A map of the workflow as it runs today, before anything is automated
  • One bounded AI-assisted system for that single process
  • Up to two integrations with tools you already pay for
  • Human approval required before any consequential action
  • Ten to twenty evaluation cases built from your real examples
  • A record of every input, action, output and failure
  • A staff walkthrough and a written handoff
  • A known-limitations sheet, and shutdown instructions

50% to begin, 50% at handoff. Third-party API and software costs are billed to you directly, not marked up.

Agent Reliability Audit

$4,500 founding rate

For a team already running an AI system that nobody has tried hard to break.

  • Where the system states things its evidence does not support
  • Prompt-injection paths through any untrusted input it reads
  • Tool permissions wider than the task requires
  • What happens on failure, and whether anyone is told
  • Whether your checkers are as independent as their number implies
  • A prioritised fix list, and one safeguard implemented

Fixed price. Findings are yours; nothing is published without your written agreement.

AI Workbench Clinic

$300 per session

One person, one working session, one system they use every week made materially better.

  • A working session on your actual files and your actual process
  • A workflow map you keep
  • Configured templates and instructions
  • One follow-up adjustment within two weeks

Paid in advance. The most common route in, and often how an organisation finds the workflow worth a sprint.

Managed Stewardship

$500 to $1,000 per month

After a system is installed: watching it, correcting it, and telling you when it drifts.

  • Monitoring of the evaluation set against live behaviour
  • A monthly report with what changed and what degraded
  • A fixed allowance of improvement work
  • Model and cost routing kept current as prices move

Only offered after a sprint or audit. There is nothing to steward before that.

How a sprint actually runs

  1. Map the workWhat happens now, who touches it, where it breaks, and what a good result looks like.
  2. Establish the baselineHow long it takes and how often it goes wrong today. Without this, no improvement can be claimed.
  3. Build the smallest useful systemOne process. Not a platform. Small enough to finish and to understand.
  4. Test against real casesYour examples, not invented ones. Including the awkward ones you would rather not send.
  5. Train, document, measureYour team runs it. The limitations are written down. The baseline is measured again.

Stage two is the one most engagements skip, and skipping it is why so many AI projects cannot say whether they worked. If we do not measure the before, there is no after to compare it to.

What I will not build

Named up front, because a boundary that only appears once you have paid is not a boundary. I do not install systems that take these actions without a person in the loop:

And four shapes I will not build at any price

The list above is about actions a system must not take on its own. These are different: they are system shapes, and no amount of review bolted on afterwards makes them safe, because the review is the part they remove.

Systems that modify or extend themselves
No self-editing prompts, self-rewriting tools, or a system that changes its own instructions between runs. Every version a system runs is one a person approved.
AI that builds or configures other AI unattended
A model may draft a config or a script. A person reads it and installs it. Nothing generates a running system and puts it into service without that step.
Improvement loops with no human in the cycle
Evaluate, revise, redeploy, repeat is the shape I refuse most firmly. It removes the only reviewer at the exact point the system starts changing fastest.
Agents that spawn or direct other agents
Orchestration that a person cannot read as a single flow is orchestration nobody can audit. One flow, one owner, one place it stops.

This is not caution borrowed from a policy document. It follows from the finding this practice is built on: a checker cannot verify past its own competence, so a system cannot safely improve past it either. Anyone selling you a self-improving loop is selling you the part where nobody is watching. If that is the engagement you want, I am the wrong person, and I would rather say so on the price page than in week two.

There is also a limit on me. I have not run a client engagement of this kind before: the research programme behind it is five months old and has published seven of its own refuted hypotheses, but the consulting practice starts with you. That is what the founding rate is for, and it is why the first two engagements are priced below what the work is worth.

Bring one workflow.

Not a strategy conversation. One process your team repeats, where it currently breaks, and what a good result would change. That is enough for me to tell you whether a sprint fits, and to say so plainly if it does not.

Three lines is a complete brief

  1. The process we repeat is…
  2. It currently breaks when…
  3. If it worked, we would…

futurisminstitute@gmail.com

That address is a plain mailbox rather than a form, and it is the same one the rest of the work uses. A dedicated one at this domain is on the list; publishing it before it can receive mail would have been the first broken promise of the engagement.

I reply to every message that describes a real process, including to say no. If a sprint is the wrong shape for what you need, I would rather tell you in week zero than in week two.