An agentic workforce is built the way a human team is built: one role at a time, with the role written down before the tool is chosen, and a clear rule about what does not go out until a person has approved it. This guide describes the method we use at Équipage IA, on our own nine roles, and then install for clients. You will find the order in which we opened ours, what we measure, and a checklist you can take as it stands.
Do you start with the model or the role definition?
With the role definition, always. The model is a technical choice you can revisit in ten minutes; an agent's remit is a management decision that touches your customer relationships. Our rule has been written in our decision log since 18 September 2026: no definition, no role, and an agent whose definition is out of date is paused.
Our definitions all use the same headings, because they are the ones you will have to answer to the day something goes wrong: the mission in one sentence; the permitted tools as a closed list, anything not on it being forbidden; what it does alone; what goes through a human; where it stops, the cases it escalates instead of settling; the automatic checks on its output; its cadence, its budget and the figures it is judged on.
The document does two jobs: it frames the agent, and it becomes the text you reread when you decide whether to give it more. In our case a definition is changed on a branch of the repository and waits for the founder's approval: an agent does not rewrite its own role.
Which role do you open first?
One, on a written task somebody already does every week and finds tedious. Not the most complex task, not the one that looks most valuable on paper: the one whose steps you know by heart, because you will know instantly whether the output is good.
Three conditions get checked before the role opens. The task arrives in writing. Answering it means fetching information from somewhere else: a catalogue, a history, a spreadsheet. And a person can check the result in seconds. If any of the three is missing, the agent will produce work nobody can verify, and checking it will cost you more than doing it. The full sort is on our AI agents for small and medium businesses page.
The opposite temptation comes up often in meetings: open five roles at once and see what happens. We turn those projects down, and not out of commercial caution. Five agents launched together means five blurred remits, no result you can attribute to anything, and nobody watching.
Where does the agent stop?
Where the company commits itself. The agent prepares, a person approves: a price, a deadline, a message to a customer, an accounting entry. That boundary is not a beginner's precaution you remove after six months, it is how the system stays usable. Write three lists for every role, and put them where the team can see them.
| The list | What goes on it | Our own example |
|---|---|---|
| What it does alone | Reversible, internal actions with no outside recipient | Read an enquiry, file it, find the reference, prepare the proposal with its sources |
| What goes through a human | Anything that goes out, anything that quotes a number, anything that promises | A message to a third party, a publication, a price, putting a page live |
| What it never does | Actions no approval makes acceptable | Print a credential, invent a figure, restart a task a human has paused |
Our agents have a fourth rule, the one that has served us best: "I don't know" and "not measured" are valid answers. An agent allowed not to know does not fill the gaps. It is also what makes checking possible: our automatic checks reject a text that states a figure without its source, and the rejection is logged.
What does an agentic workforce need in common?
The same three things a human team needs: shared rules, shared tools, and a memory of decisions.
A charter. Ours runs to a few pages and every agent rereads it at the start of every task. It says how you ask a colleague for something, what a request and a finished piece of work contain, and what is not allowed out. It sets two limits that prevent loops: at most three levels of cascading requests, and at most two round trips between two agents on the same subject, after which a person settles it.
Shared tools, named. One command per use, capped, and a table saying which. We learned that rule the hard way: an agent declared a tool "not configured" when it was working, because it had gone looking for a key in the files instead of using the command provided. Since then the instruction is written down: before declaring a tool unavailable, run the command in the table and quote the exact error.
A decision log. Every call is dated, written down, filed in one place. It is what lets an agent woken up three weeks later know why things are the way they are, and lets you revisit a decision without replaying the discussion. A conversation nobody wrote down is gone.
A board the work moves across: one card per task, showing who produced what, from which source, and what it cost. That is where you approve, and where you notice an agent going in circles.
How do you know it is working?
By measuring three things, from the first week, and nothing more than three at the start.
- Cost per task. Not the model's price per million tokens: the full cost of one finished task, the only figure comparable to the time it used to take you. In our setup every role has a daily ceiling and the command refuses the spend before making it. We wrote a whole article about it: pay for the task, not the token.
- Time given back. How many minutes the person who used to do the task no longer spends on it, and what they do instead. Measure it before you start, or you never will.
- Errors and rework. How many outputs the automatic checks rejected, how many you sent back, and why. A reason that comes up three times is not an agent mistake: it is a rule missing from its definition.
Two traps have cost us time. The first: changing the model to fix a mediocre result, when the problem was the information the agent had been given (tidy up what the agent needs to know). The second: mistaking slowness for incapacity, when the wait came from the way the steps were chained together (why your AI agent is slow).
In what order do you open the roles, and why?
In the order of what gets you closer to a customer, not the order of what is easy. The principle: every role you open must either produce something that sells, or make the next role possible. The table below is the sequence we set on 18 September 2026 in our decision log, with the trigger chosen for each role. It is a decision about sequence, not a report: a role opens on the day its trigger is ticked, and the ninth, customer relations, is still waiting for its own: the first paying customer.
| Order | Role | Trigger, and why at that point |
|---|---|---|
| 1 | Prospecting | The role that gets us closer to a customer. It fixes the building blocks reused afterwards: code for what must be reliable, a model for what must be written, a check before approval, a tracked budget |
| 2 | Market watch | It feeds everything else with verified subjects and touches no third party: a risk-free role to break the system in |
| 3 | Operations | Only useful once there are two roles to coordinate |
| 4 | Development | As soon as there is write access to the repository: this is the role that equips the others |
| 5 | Social media | As soon as the account is open and the voice approved |
| 6 | Contact follow-up | As soon as the first prospecting messages go out: before that there is nothing to follow |
| 7 | Search and writing | As soon as the website can take pages. An article with nowhere to publish is worthless |
| 8 | Studio | As soon as the design system is approved, otherwise you produce visuals you will redo |
| 9 | Customer relations | At the first paying customer, not before |
Three transferable lessons. Every role waits for a trigger you can tick, never an urge. A definition gets revised after opening: our writing role, designed as an executor, became responsible for its own field because the person choosing the subjects was the bottleneck. And a role can stay in break-in mode for a long time, with its scheduled tasks paused: pausing is not a failure, it is the default.
The ten-point checklist
To go through before you open your first role. Ten written answers; one page is enough.
- The target task arrives in writing, it repeats, and somebody already does it every week.
- You have measured how long it takes today, before touching anything.
- The role definition is written: mission, permitted tools as a closed list, what it does alone, what goes through a human, where it stops.
- Access rights are listed one by one, with the use of each. Any access nobody can explain is removed.
- The three lists are posted: alone, approved, never.
- A named person approves, and knows how many minutes a day that will take them.
- Every output leaves a trace you can consult: what was read, prepared, sent, rejected, and on what basis.
- A spending ceiling is set, and the system refuses the spend before making it.
- You know how to stop the agent, somebody on your side knows how to do it, and you know what happens to the work in progress.
- Three figures are tracked weekly: cost per task, time given back, errors and rework. A rework reason that appears three times becomes a rule in the definition.
If you can tick all ten, open the role. If one is missing, that is the one to deal with before writing a single line of configuration.
What if the right answer were to build nothing?
Sometimes it is. A task that changes every month, a process nobody has written down, data too dirty to use: in those cases an agent amplifies the mess instead of reducing it. It also happens that fixed-rule automation, with no artificial intelligence at all, is enough and costs less to maintain. That is one of the three possible conclusions of our discovery audit. Before connecting an agent to your tools, read the four questions to ask as well.
FAQ
How many roles make it a team?
Two is enough, as soon as they hand work to each other and a shared framework says how. What makes a team is not the count: it is that the work moves in a traceable way and a person keeps the decision. Start with one role, add the second when the first runs without waking you up at night.
Do you need technical skills in-house?
To steer, no: you need someone who knows the trade, can say what good looks like, and accepts approving every day. To install and maintain, yes, in-house or with a supplier. The role that cannot be delegated is the approving one.
How much of the owner's day does this take?
A short but daily slot to approve what is waiting. In our case the founder's approval time is capped at one hour a day, everything included, and that constraint is what limits how many roles are open at once. If your system produces more than you can read, it is useless.
What happens when an agent gets it wrong?
Whatever you planned, or nothing good. In our setup a rejected output is logged with its reason, goes back to its author, and a reason that appears three times becomes a written rule in the role definition. Mistakes do not disappear: they turn into rules. What matters is that none can reach a customer without a person having seen it go past.
